RL Practitioner Postmortem: Alternating SFT, Souping and RL Yields a Good Explore-Exploit Cycle
tensorqt · x · 2026-10-07
Trainer @tensorqt reflects on an RL run: in hindsight they could have done more SFT. The training curve looks messy partly because it is, but also because they found that alternating SFT - souping - RL gives a good exploration and exploitation cycle.
A first-hand recipe note from a practitioner on composing post-training stages, useful for teams doing RL fine-tuning.
More from Research
- Thinking Machines explains why LLM inference stays nondeterministic even at temperature 0 — JFPuget · 2026-10-07
- QUEEN paper distills AlphaZero-style chess search into language for LLMs — danqi_chen · 2026-10-07
- COLM paper asks whether VLMs can internalize tool calls in latent space instead of calling them — PMinervini · 2026-10-07
- CLeaR 2027 Opens Its Submission Portal — ArthurGretton · 2026-10-07
- MIT paper: a minimalist agent loop that passes history as code variables beats Letta and ACE at half the cost — rohanpaul_ai · 2026-10-07
- doodlestein releases explainer video alongside open-source companion repo — doodlestein · 2026-10-07