SFT Shifts Reasoning Language, RL Fixes Formatting in Low-Resource SFT
KIEFERSA · hf · 2026-08-21
Fine-tuning MoE models on low-resource languages shifts reasoning into that language without harming accuracy. RL with verifiable rewards fixes formatting and leakage defects.
More from Research
- Eris theorem proving environment updates with direct manipulation — round · 2026-08-21
- Nature Paper: Consumer AI shifts from info tool to healthcare pathway control — EricTopol · 2026-08-21
- DeepMind uses AlphaEvolve to improve matrix multiplication bound — rohanpaul_ai · 2026-08-21
- Hydra-0 unifies action language for physics simulation and robot control — YunzhuLiYZ · 2026-08-21
- David Baker's Team Publishes on Accelerating AI Protein Design Validation Loop — AllThingsApx · 2026-08-21
- CP-MoE: Freeze the Base, Tune Just 1.5% Params, No Forgetting — flosalim · 2026-08-21