Mixed SFT Outperforms Next-Chunk RL Without CoT Data

Yinhao Tang · hf · 2026-08-27

Revisiting training strategies without CoT data, the study finds that mixed supervised fine-tuning on combined reasoning corpora outperforms next-chunk reinforcement learning in both efficiency and final accuracy across mathematical and out-of-domain tasks.

Original post →

More from Research

Research channel →