Princeton team's Coverage Principle paper explains why xent SFT underperforms for RL, proposes TailSFT

canondetortugas · x · 2026-09-12

An arXiv paper by Fan Chen, Sadhika Malladi, Dylan Foster et al. offers a theory of how pre-training enables post-training via coverage — the probability mass a model puts on high-quality responses. Coverage is necessary and sufficient for Best-of-N style post-training to work and predicts downstream performance better than cross-entropy, avoiding spurious length dependence. Follow-up work TailSFT replaces standard xent SFT with a lightweight, principled method that directly improves coverage and boosts post-RL performance.

Related event: DeepMind and Princeton Propose TailSFT to Boost RL Performance(3 posts)→

Original post →

More from Research

Research channel →