Microsoft TailSFT boosts code pass@16 by up to 16.8 points
dair_ai · x · 2026-08-31
Microsoft researchers propose TailSFT, a method to improve the standard SFT pipeline for better Reinforcement Learning performance.
- Problem: Standard SFT continues spending gradient on sequences the model has already fit, narrowing the distribution RL needs to explore later.
- Method: TailSFT filters out these sequences during training, focusing learning on the under-modeled tail of the data.
- Results: On OLMo-3 7B, pass@16 improves by up to 16.8 points absolute on coding and 3.1 on math. These higher-coverage checkpoints lift final pass@1 after GRPO by up to 3.9 points, and in some settings, early reward climbs 2.5x faster than matched standard SFT runs.
More from Research
- LiteMol-1 generates drug candidates on M1 Max in 30 seconds — CatAstro_Piyush · 2026-09-01
- Google Paper: Autonomous AI Research Hallucinates 90% Without Checks — rohanpaul_ai · 2026-09-01
- RLHF impact on tokens: unconscious shifts vs conscious choices — voooooogel · 2026-09-01
- On token layers and consciousness in RLHF — voooooogel · 2026-09-01
- CommerceAgentBench released: Qwen leads open-weight models — Alibaba_Qwen · 2026-09-01
- Discussion on Why Universal Time Series Models Work — Afinetheorem · 2026-09-01