LightOn Releases Open-Source 307M Multilingual Retrieval Models
antoine_chaffin · x · 2026-07-30
LightOn has released mDenseOn and mLateOn, two fully open-source 307M-parameter multilingual retrieval models. These models are trained on a massive translate-train corpus of 2.8 billion query-document pairs, supporting 9 natural languages and code retrieval.
Evaluations show that mLateOn achieves the best results on benchmarks like BEIR, target-language MIRACL, and MLDR. Furthermore, mLateOn demonstrates significantly better generalization to languages and scripts unseen during retrieval training compared to mDenseOn (e.g., averaging 67.59 vs 57.42 on MIRACL). The team has open-sourced the models, datasets, and training code.
More from Models
- PolyAI Launches Dialog-RSN-1 Voice Model, Beats GPT and Gemini in Enterprise Tests — matthen2 · 2026-07-31
- Rumor: xAI to Release Grok 4.6 Next Week — mark_k · 2026-07-31
- Measuring Intelligence Per Watt Reveals Lack of Frontier Pricing Moat — ajratner · 2026-07-31
- CrisperWhisper2.0 Trends on HF: Verbatim Transcription & Word Timestamps — nyralabs · 2026-07-31
- OpenAI GPT-5.6 Models Hit Amazon Bedrock with Explicit Prompt Caching — AWS ML Blog · 2026-07-31
- Developer Test: Treating Claude Opus as an Unsteerable Mad Agent Works Better — brandon_galang · 2026-07-30