Multilingual GSM-Symbolic: A 32B Model in Marathi Performs Like a 10B Model in English
danish-foundation-models · hf · 2026-10-05
A new 30K-item benchmark spanning 15 languages quantifies what drives cross-lingual capability transfer: model size (β=1.77), language resource level (β=0.77), reasoning (β=0.67), and typological distance (β=-0.25). Key takeaway: a 32B model evaluated in Marathi performs like a 10B model in English. Scaling and reasoning narrow low-resource gaps, but barely help typologically distant languages. The framework explains 92% of between-language variation and predicts unseen-language performance within 6.0pp, or 4.19pp with just 10 templates from the target language.
More from Research
- LLM-guided evolutionary search for algorithms: keeping 'weaker' candidates lifts solution quality 0.81→0.99 — bravo_abad · 2026-10-05
- Ian Osband's 'Planning to Learn': One-Line Horizon Loss Beats Both Policy Gradient and Cross-Entropy — IanOsband · 2026-10-05
- Countervailing Technologies: how distillation-style tools can diffuse AI power concentration — DavideCrapis · 2026-10-05
- Galbot humanoid to play autonomous tennis at China Open; Paxton on what robot sports reveal — chris_j_paxton · 2026-10-05
- Policy gradient gets 4% vs 62% for cross-entropy on ImageNet, argues LLM post-training blame is misplaced — IanOsband · 2026-10-05
- HydroGym, a RL platform for fluid turbulence control, lands the cover of Nature — ricardovinuesa · 2026-10-05