TailSFT Paper: Filtering Fitted Sequences in SFT Boosts Post-RL pass@1 by up to 4%
rohanpaul_ai · x · 2026-09-04
- arXiv paper "TailSFT: Filtered Fine-Tuning Improves Post-Training Performance" (Sadhika Malladi, Samy Jelassi, et al., Microsoft) argues existing SFT pipelines may not yield checkpoints best suited for RL, building on coverage/pass@K as predictors of post-RL performance.
- Method: filter out already-fit sequences during SFT, focusing learning on the under-modeled tail of the data distribution; design choices validated via controlled experiments and theory.
- Results: on OLMo-3 7B, pass@16 on math/coding improves up to 17% absolute with minimal overhead, and subsequent GRPO runs gain up to 4% absolute pass@1. Includes a lightweight diagnostic for when TailSFT helps, advocating stage-aware checkpoint selection.
Related event: Microsoft's TailSFT: Filtered Fine-Tuning Improves RL Performance(2 posts)→
More from Research
- New blog derives Curiosity-Driven Tree Search as a scalable AlphaZero alternative — CatAstro_Piyush · 2026-09-04
- IP-Adapter at 0.6 overrides prompts; lower it and 4-view character consistency breaks — God_Speedmyboy · 2026-09-04
- 7M-Parameter Tiny Recursion Model Hits 45% on ARC-AGI-1 — CatAstro_Piyush · 2026-09-04
- New Paper Tunes Training-Time MSA Depth Distribution to Beat AlphaFold2/3 — MoAlQuraishi · 2026-09-04
- KV cache doesn't enable information passing, it just saves compute — scaling01 · 2026-09-04
- YC-affiliated team offers $500k and free YAM arms for dexterous manipulation benchmarks — ycombinator · 2026-09-04