TailSFT: skip already-learned SFT examples to boost post-RL pass@k
canondetortugas · x · 2026-09-12
- DeepMind researchers propose TailSFT, a lightweight SFT tweak that improves coverage and post-RL performance.
- Motivation: RL is expensive; prior work (1yr ago) showed vanilla cross-entropy SFT isn't the best RL prep.
- Method: record each example's pre-SFT loss L0; during SFT, exclude the examples most-improved-relative-to-L0 from backprop (sort-and-slice or zero-masking), since the model has already learned them.
- Result: better pass@k starting point for k>1, better post-RL results. Lucas Beyer calls it simple and intuitive.
Related event: DeepMind and Princeton Propose TailSFT to Boost RL Performance(3 posts)→
More from Models
- Claude 5x Pro User Claims Post-Reset Quota Burns Far Faster Than Day One — unknown9645 · 2026-09-12
- GPT-6 Astra shows 'step change' in spatial reasoning, solving 7/100 robot tasks vs zero for rival — The Decoder · 2026-09-12
- EMNLP paper: Prompt2Box uncovers entailment structure to find LLM weaknesses — windx0303 · 2026-09-12
- DeepSeek V4.1 Flash lands on Merge Gateway, tops Terminal-Bench 2.1 at 90.6 — shensi · 2026-09-12
- 2019 Pruning Experiment Cited to Claim 96% of GPT-5's Weights Are Useless — TinfoilTricorn · 2026-09-12
- DeepSeek V4-Pro-0813 and GLM-5.3 Launch on Nebius, Targeting Coding Agents — Arindam_1729 · 2026-09-12