Princeton team's Coverage Principle paper explains why xent SFT underperforms for RL, proposes TailSFT
canondetortugas · x · 2026-09-12
An arXiv paper by Fan Chen, Sadhika Malladi, Dylan Foster et al. offers a theory of how pre-training enables post-training via coverage — the probability mass a model puts on high-quality responses. Coverage is necessary and sufficient for Best-of-N style post-training to work and predicts downstream performance better than cross-entropy, avoiding spurious length dependence. Follow-up work TailSFT replaces standard xent SFT with a lightweight, principled method that directly improves coverage and boosts post-RL performance.
Related event: DeepMind and Princeton Propose TailSFT to Boost RL Performance(3 posts)→
More from Research
- Pretraining on Europe alone beats global data across every task, 10-21 point gap in remote sensing study — anselm · 2026-09-12
- New CET method traces how psychological constructs emerge across LLM layers — GolinoHudson · 2026-09-12
- Digital fly brain lets researchers simulate lesions and study behavior changes — BraydonDymm · 2026-09-12
- Digital fly brain with 165,122 traced neurons lets you simulate lesions and watch behavior change — BraydonDymm · 2026-09-12
- Three runs, two significant losses: your cheap-model switch may just have been luck — Ok-Challenge-7810 · 2026-09-12
- TailSFT: skip already-learned SFT examples to boost post-RL pass@k — canondetortugas · 2026-09-12