LightOn Releases Multilingual Retrieval Models and 2.8B Training Pairs
IgorCarron · x · 2026-07-30
LightOn has introduced two new multilingual retrieval models, mDenseOn and mLateOn, focusing on multilingual, long-context, and code retrieval.
The highlight of this release is its full openness. Alongside the models, the team has open-sourced a massive training corpus comprising 2.8 billion curated query-document pairs across 9 languages, including French, Italian, German, and Arabic. The developers claim these models achieve frontier performance in multilingual retrieval.
More from Models
- Report: Anthropic Opus 5 Incoming, Kimi K3 to Be Largest OSS Model — WolframRvnwlf · 2026-07-30
- Google Launches Gemini Robotics ER 2 with Enhanced Multi-Robot Collaboration — OfficialLoganK · 2026-07-30
- Google Launches Gemini Robotics ER 2 for Embodied Reasoning — OfficialLoganK · 2026-07-30
- Anthropic Lacks Standard Completions Endpoint; Fair Model Eval Needs Unified Harness — altryne · 2026-07-30
- Polymarket Odds: 54% Chance of New Gemini Pro Release by Next Month — Polymarket · 2026-07-30
- P-Image-Ideogram Hits Pareto Frontier for Image Gen Speed and Cost — _akhaliq · 2026-07-30