TailSFT: tail-weighted SFT improves coverage and post-RL performance, says new paper

canondetortugas · x · 2026-09-12

Sadhika Malladi's team proposes TailSFT: since RL is expensive and every step should count, and following their earlier finding that xent SFT isn't the best preparation for RL, TailSFT offers a lightweight, principled way to directly improve coverage and achieve better post-RL performance. Built on OLMo and praised by researchers like finbarrtimbers.

Related event: DeepMind and Princeton Propose TailSFT to Boost RL Performance(3 posts)→

Original post →

More from Research

Research channel →