TailSFT: Microsoft Paper Shows SFT Can Wreck RL Starting Points, Gains up to 4% pass@1
rohanpaul_ai · x · 2026-09-04
- A new Microsoft + UC San Diego paper shows standard SFT can make a model look better on evals while actually being a worse starting point for RL, by wiping out rare correct behaviors RL needs to reinforce.
- TailSFT: filter out training sequences whose loss has already dropped most relative to the base model, shifting learning toward under-modeled regions of the data distribution so correct responses stay reachable under repeated sampling.
- Results: on OLMo-3 7B, pass@16 on math and coding improves up to 17% absolute with minimal overhead; subsequent GRPO runs see up to 4% absolute pass@1 gains (up to 3.93 points). The paper also introduces a lightweight diagnostic for when TailSFT helps.
Related event: Microsoft's TailSFT: Filtered Fine-Tuning Improves RL Performance(2 posts)→
More from Research
- Reusing reasoning traces as pre-context state lifts long-context accuracy in 26 of 27 tests — burny_tech · 2026-09-04
- Running Full Scientific Experiments on Robots with Claude: New Protocol in Under 10 Minutes — yawnxyz · 2026-09-04
- Deep Learning for RNA Design Makes Science Cover, AI Matches Expert Humans on Pseudoknots — rishabh16_ · 2026-09-04
- 753B model 'thinks', 4B model writes: latent-space handoff claimed to be 20x faster — burny_tech · 2026-09-04
- Vals AI launches SRE-Bench, a cybersecurity benchmark testing LLMs on binary reverse engineering — dyn___ · 2026-09-04
- Emergent Misalignment Is Predictable Generalization, Not a Magic "Evil Persona" — burny_tech · 2026-09-04