LightOn Releases Open Multilingual Retrieval Models with 2.8B Training Pairs
antoine_chaffin · x · 2026-07-30
LightOn has released mDenseOn and mLateOn, 307M-parameter multilingual retrieval models.
- Performance: Achieves state-of-the-art on BEIR and MLDR, while remaining highly competitive on MIRACL and MTEB Code.
- Training Data: Translated validated English pairs into 8 languages, creating a massive 2.8 billion pair cross-lingual corpus.
- Generalization: mLateOn demonstrates strong zero-shot capabilities, remaining competitive on 13 languages unseen during retrieval training (e.g., Chinese, Japanese, Korean).
- Fully Open: Models, datasets, and training code are open-sourced under Apache 2.0.
More from Models
- How Kimi K3 engineered its way to the frontier: Delta Attention, expert balancing, AgentENV — noninertialframe96 · 2026-07-31
- Rumor: xAI to Release Grok 4.6 Next Week — mark_k · 2026-07-31
- Measuring Intelligence Per Watt Reveals Lack of Frontier Pricing Moat — ajratner · 2026-07-31
- CrisperWhisper2.0 Trends on HF: Verbatim Transcription & Word Timestamps — nyralabs · 2026-07-31
- OpenAI GPT-5.6 Models Hit Amazon Bedrock with Explicit Prompt Caching — AWS ML Blog · 2026-07-31
- Developer Test: Treating Claude Opus as an Unsteerable Mad Agent Works Better — brandon_galang · 2026-07-30