Apple's large-scale study: GRPO reasoning in native languages nearly matches English
Apple ML Research · rss · 2026-08-18
Apple ML Research publishes "GRPO Beyond English," a large-scale empirical study of multilingual and non-English GRPO, addressing how English-centric current RLVR research is. The study spans a wide range of base models, training languages, and different reasoning-language rewards.
Key finding: training models to reason in their native language often leaves only a small gap compared to training for English reasoning—suggesting reasoning RL can be done effectively in local languages.
More from Research
- InfinityEdit: Infinite Video Editing via Lightweight Adapter — Yunze Tong · 2026-08-24
- Tencent Benchmarks Hybrid-Thinking MLLMs for Response Alignment — tencent · 2026-08-24
- CLEAR Adapter Routing Balances LLM Safety and Utility — UIUC-CS · 2026-08-24
- Llama-Mobile: 2.7-Bit Quantization Shrinks Llama 3.2 Vision 11B to 3.7GB for Phones — Luka Ribar · 2026-08-24
- Retriever: A Framework for Asynchronous, Closed-Loop Robot Agents — ZeYanjie · 2026-08-24
- Converting GMMs ↔ PEFs for fast KLD approximation — FrnkNlsn · 2026-08-24