Fixing RL collapse: 59 SFT steps to reinject exploration 'potential energy', researcher says
tensorqt · x · 2026-10-07
Researcher tensorqt argues there's no clearly sample-efficient fix for RL mode collapse; he views it as "potential energy" the model spends during RL at the cost of exploration, and the fix is reinjecting that energy — not necessarily via SFT, though SFT is fastest: just 59 fine-tuning steps in their last RL run yielded significant gains that beat the frontier.
A replier adds that collapse usually comes with entropy collapse (fixing that may make the step unnecessary), notes pretraining souping worked surprisingly well, and doubts the critic helped at all.
More from Research
- Looped transformers are trending again, tracing roots back to a 2016 paper — natanielruizg · 2026-10-07
- Fixed token codes suffice: 1.7B LM trains without a trainable input embedding table — A. Bochkov · 2026-10-07
- EmbeddingGemma 2 hands-on: 740M multimodal embeddings for search and RAG, runnable on a free T4 — Prompt Engineering · 2026-10-07
- Isomorphic, DeepMind and Meta join DOE-NIH partnership to build an AI model of the cell — snikolov · 2026-10-07
- Bi-manual mobile UMI demo unlocked for robot manipulation data collection — neurosp1ke · 2026-10-07
- Researcher presents Meta-Harness and Combee at COLM 2026, seeks industry roles — Kangwook_Lee · 2026-10-07