Mixed SFT Outperforms Next-Chunk RL Without CoT Data
Yinhao Tang · hf · 2026-08-27
Revisiting training strategies without CoT data, the study finds that mixed supervised fine-tuning on combined reasoning corpora outperforms next-chunk reinforcement learning in both efficiency and final accuracy across mathematical and out-of-domain tasks.
More from Research
- Paper reveals why PPO value functions fail, proposes BPCO for stable training — heghbalz · 2026-08-27
- Gaussian fiddling brings facial expressions to Clug — repligate · 2026-08-27
- Lightwheel and Hugging Face release 100k-hour egocentric dataset for Physical AI — vanstriendaniel · 2026-08-27
- PyTorch Ecosystem Adds Perforated, TokenSpeed, and 8 Others — zhyncs42 · 2026-08-27
- V-Rubrics: Improving Visual Faithfulness via Rubric-Based Reinforcement Learning — liuziwei7 · 2026-08-27
- FlashKDA ref impl lower bound set to -5 months ago — YouJiacheng · 2026-08-27