LightOn Open-Sources Multilingual Retrieval Models Topping BEIR & MTEB Leaderboards
antoine_chaffin · x · 2026-07-30
LightOn has released mDenseOn and mLateOn, two fully open-source 307M-parameter multilingual retrieval models. These models achieve state-of-the-art results across English general-domain retrieval (BEIR), long-document retrieval (MLDR), multilingual retrieval (MIRACL), and code retrieval (MTEB Code).
Built upon a validated English data recipe, the pipeline was extended to eight additional languages via translate-train, creating a massive 2.8B-pair corpus. Most notably, the late-interaction model (mLateOn) demonstrates strong zero-shot generalization, remaining competitive in languages and scripts entirely absent from retrieval training (such as Cyrillic, Japanese, Korean, and Chinese). This effectively removes the main limitation of translate-train pipelines with dense models. Models, datasets, and training code are fully open-sourced.
More from Research
- Breakthrough in Nanopore Protein Sequencing Achieves 98% Accuracy — NikoMcCarty · 2026-07-30
- AI Model Improves Attack on 7-Round AES, Sponsored by Anthropic with $100k Compute — jedisct1 · 2026-07-30
- AI Detector Fail: LLM-Generated Random Numbers Flagged 100% AI — nrehiew_ · 2026-07-30
- Repeating LLM Evals Doesn't Boost Accuracy Due to Correlated Errors — randal_olson · 2026-07-30
- Study Finds Majority Voting for LLM Evals Ineffective; Human Labels Still Essential — randal_olson · 2026-07-30
- Hugging Face Launches Agentic Challenge to Advance Alzheimer's Research — _lewtun · 2026-07-30