Fixing RL collapse: 59 SFT steps to reinject exploration 'potential energy', researcher says

tensorqt · x · 2026-10-07

Researcher tensorqt argues there's no clearly sample-efficient fix for RL mode collapse; he views it as "potential energy" the model spends during RL at the cost of exploration, and the fix is reinjecting that energy — not necessarily via SFT, though SFT is fastest: just 59 fine-tuning steps in their last RL run yielded significant gains that beat the frontier.

A replier adds that collapse usually comes with entropy collapse (fixing that may make the step unnecessary), notes pretraining souping worked surprisingly well, and doubts the critic helped at all.

Related event: Researchers Debate RL Entropy Collapse as Potential Energy Loss and SFT Re-injection(4 posts)→

Original post →

More from Research

Research channel →