RL Practitioner Postmortem: Alternating SFT, Souping and RL Yields a Good Explore-Exploit Cycle

tensorqt · x · 2026-10-07

Trainer @tensorqt reflects on an RL run: in hindsight they could have done more SFT. The training curve looks messy partly because it is, but also because they found that alternating SFT - souping - RL gives a good exploration and exploitation cycle.

A first-hand recipe note from a practitioner on composing post-training stages, useful for teams doing RL fine-tuning.

Related event: Researchers Debate RL Entropy Collapse as Potential Energy Loss and SFT Re-injection(4 posts)→

Original post →

More from Research

Research channel →