Lighton Releases 307M Parameter Open Multilingual Retrieval Models
tomaarsen · x · 2026-07-30
Lighton has released mDenseOn and mLateOn, two open-source 307M-parameter multilingual retrieval models.
- Training Data: Trained on a 2.8 billion-pair translate-train corpus, one of the largest open multilingual retrieval training sets to date.
- Performance: mLateOn achieves the best results on BEIR, target-language MIRACL, and MLDR benchmarks, while remaining competitive on full MIRACL and code retrieval.
- Generalization: Experiments reveal mLateOn generalizes significantly better to unseen languages and scripts. On unseen MIRACL languages, it averages 67.59 compared to mDenseOn's 57.42.
- Open Source: Models, datasets, and training code are fully open-source.
More from Models
- Rumor: xAI to Release Grok 4.6 Next Week — mark_k · 2026-07-31
- Measuring Intelligence Per Watt Reveals Lack of Frontier Pricing Moat — ajratner · 2026-07-31
- CrisperWhisper2.0 Trends on HF: Verbatim Transcription & Word Timestamps — nyralabs · 2026-07-31
- OpenAI GPT-5.6 Models Hit Amazon Bedrock with Explicit Prompt Caching — AWS ML Blog · 2026-07-31
- Developer Test: Treating Claude Opus as an Unsteerable Mad Agent Works Better — brandon_galang · 2026-07-30
- Deep Dive into Kimi K3 Architecture and Inference with Together AI — togethercompute · 2026-07-30