TailSFT: Skipping SFT Gains Improves pass@16 and RL Exploration in GRPO
tw_killian · x · 2026-09-15
A shared research finding argues that RL (especially GRPO) depends on exploration, and conventional SFT can hurt it: fine-tuning oversamples tasks the model already does well, sharpening behavior and limiting exploration.
- TailSFT records each example's pre-SFT loss L0, then excludes the most-improved examples from backprop during SFT
- Result: lower pass@1 than standard SFT but higher pass@16, with diversity translating into better RL training dynamics
- Described as a simple, intuitive trick for giving RL a better pass@k starting point
More from Research
- ORQA paper tests LLM knowledge across 116 occupations; top models score just ~60% — soumitrashukla9 · 2026-09-15
- Hand-Deriving the VAE in 11 Steps: One Diagram Teaches KL Divergence and Diffusion Loss — ProfTomYeh · 2026-09-15
- Fly connectome reveals fast-weight continual learning neurons — a skill current LLMs lack — tszzl · 2026-09-15
- Hyperstition claims 62% pretraining cost cut and 1.7B math model beating Qwen3 — nick_linck · 2026-09-15
- Biologist Eörs Szathmáry warns AGI's replication speed makes it a runaway biosphere risk — danfaggella · 2026-09-15
- Survey: Asian AI researchers worry more about AI risks than Western peers — KatjaGrace · 2026-09-15