307M Parameter Model mLateOn Excels in Multilingual, Long Doc, and Code Retrieval
antoine_chaffin · x · 2026-07-30
The author introduces a new 307M parameter retrieval model that delivers strong performance across English, multilingual, long-document, and code retrieval tasks.
Key Highlights & Ablations:
- Code Retrieval: Despite no code-specific pre-training, fine-tuning on the LateOn-Code dataset yields strong scores of 71.5 (mDenseOn) and 73.5 (mLateOn) on the MTEB Code benchmark.
- Data Efficiency: Ablation studies reveal that the model only requires a small amount of English/cross-lingual data to achieve robust generalization.
More from Research
- Breakthrough in Nanopore Protein Sequencing Achieves 98% Accuracy — NikoMcCarty · 2026-07-30
- AI Detector Fail: LLM-Generated Random Numbers Flagged 100% AI — nrehiew_ · 2026-07-30
- Repeating LLM Evals Doesn't Boost Accuracy Due to Correlated Errors — randal_olson · 2026-07-30
- Hugging Face Launches Agentic Challenge to Advance Alzheimer's Research — _lewtun · 2026-07-30
- Papers with Code Adds Artifact Filters; NVIDIA Tops Org Contribution Graphs — NielsRogge · 2026-07-30
- Cell Paper: Brute-Force Scaling Won't Solve Biological Modeling — anshulkundaje · 2026-07-30