Cohere Labs paper: data mixing lifts in-language reasoning above 93% across 60 languages
ShahabBakht · x · 2026-09-24
Cohere Labs research scholar Mehrnaz Mofakhami will present her new paper, Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning (arXiv:2609.10445), in a live session on Sept 28.
Paper highlights
- Problem: reasoning models are overwhelmingly English-centric — they reason in English regardless of prompt language, hurting non-English users and discarding target-language knowledge
- Proposes L2 reasoning: models reason consistently in the user's language, approached via data-centric optimization of SFT data composition and scheduling
- Built Tiny Aya L2-Thinker at 3.35B, achieving an L2 reasoning rate above 93% across 60 languages on 6 benchmarks spanning math, commonsense, instruction following, open-ended generation and cultural reasoning, while keeping performance
More from Research
- Neel Nanda's team launches WorkspaceBench, an eval to test interpretability tools — burny_tech · 2026-09-24
- SchrödingerRepo: rewriting repos exposes LLM memorization on SWE-bench — SJTU · 2026-09-24
- Survey: memory mechanisms for autoregressive video generation — Harold Haodong Chen · 2026-09-24
- Robotics' real frontier is post-training: self-play in sim closes the demo-to-deployment gap — ZGojcic · 2026-09-24
- Pedro Domingos: citation counts that don't divide by author count are unserious — lemire · 2026-09-24
- Is RLHF the artificial selection of LLMs? A Darwinian analogy for model training — Short-Search-4892 · 2026-09-24