Multilingual GSM-Symbolic: A 32B Model in Marathi Performs Like a 10B Model in English

danish-foundation-models · hf · 2026-10-05

A new 30K-item benchmark spanning 15 languages quantifies what drives cross-lingual capability transfer: model size (β=1.77), language resource level (β=0.77), reasoning (β=0.67), and typological distance (β=-0.25). Key takeaway: a 32B model evaluated in Marathi performs like a 10B model in English. Scaling and reasoning narrow low-resource gaps, but barely help typologically distant languages. The framework explains 92% of between-language variation and predicts unseen-language performance within 6.0pp, or 4.19pp with just 10 templates from the target language.

Original post →

More from Research

Research channel →