SFT Shifts Reasoning Language, RL Fixes Formatting in Low-Resource SFT

KIEFERSA · hf · 2026-08-21

Fine-tuning MoE models on low-resource languages shifts reasoning into that language without harming accuracy. RL with verifiable rewards fixes formatting and leakage defects.

Original post →

More from Research

Research channel →