RL isn't just sharpening existing skills — pre-LLM RL trained from random init
cephaloform · x · 2026-09-26
The author pushes back on the popular claim that RL can only sharpen capabilities already latent in a model, noting it falls apart when you recall that nearly all pre-LLM-era RL was done on random-init models. He quips that the idea of a random init containing all possible capabilities is "almost poetic."
More from Research
- Lawrence Krauss podcast asks whether AI will supercharge scientific paper mills — willcb · 2026-09-26
- Framework-free prototype learner lets local LLMs learn corrections instantly, no fine-tuning — thisdudelikesAI · 2026-09-26
- Building an eval harness for ChatGPT, and the contamination problem of seen solutions — tak3sh8 · 2026-09-26
- Maximize intelligence per flop: decay adaptivity to trade for precision — willcb · 2026-09-26
- Why specialized agents that update their own weights beat generalist models — willcb · 2026-09-26
- Project Callisto: Nine Years and 500 Samples Find No Evidence of Cold Fusion — skdh · 2026-09-26