TailSFT: Skipping SFT Gains Improves pass@16 and RL Exploration in GRPO

tw_killian · x · 2026-09-15

A shared research finding argues that RL (especially GRPO) depends on exploration, and conventional SFT can hurt it: fine-tuning oversamples tasks the model already does well, sharpening behavior and limiting exploration.

Original post →

More from Research

Research channel →