TailSFT: tail-weighted SFT improves coverage and post-RL performance, says new paper
canondetortugas · x · 2026-09-12
Sadhika Malladi's team proposes TailSFT: since RL is expensive and every step should count, and following their earlier finding that xent SFT isn't the best preparation for RL, TailSFT offers a lightweight, principled way to directly improve coverage and achieve better post-RL performance. Built on OLMo and praised by researchers like finbarrtimbers.
Related event: DeepMind and Princeton Propose TailSFT to Boost RL Performance(3 posts)→
More from Research
- Whole-Brain Fly Simulation Hits 91% Behavior Accuracy With No Training — alexcovo_eth · 2026-09-12
- AI bio capability hype vs wet lab reality: models fail over half the time — nlarusstone · 2026-09-12
- smolbenchmark ranks sub-8GB models by speed, tok/J and heat on your own hardware — East-Muffin-6472 · 2026-09-12
- Clay Math Institute says the Navier-Stokes problem 'has apparently been settled' — badumtsssst · 2026-09-12
- Decagon shares 19+ ablations on using GEPA for test-driven prompt optimization in production — kastnerkyle · 2026-09-12
- GraphED: graph-based AI learns how solids deform by sharing law structure across materials — bravo_abad · 2026-09-12