Apple's large-scale study: GRPO reasoning in native languages nearly matches English

Apple ML Research · rss · 2026-08-18

Apple ML Research publishes "GRPO Beyond English," a large-scale empirical study of multilingual and non-English GRPO, addressing how English-centric current RLVR research is. The study spans a wide range of base models, training languages, and different reasoning-language rewards.

Key finding: training models to reason in their native language often leaves only a small gap compared to training for English reasoning—suggesting reasoning RL can be done effectively in local languages.

Original post →

More from Research

Research channel →