Microsoft's TailSFT: Filtered Fine-Tuning Improves RL Performance

A Microsoft and UCSD paper proposes TailSFT, which filters already-fitted samples before SFT to preserve rare correct behaviors. The approach improves post-RL pass@1 by up to 3.93 percentage points.

2026-09-04 ~ 2026-09-04 · 2 related posts