RL isn't just sharpening existing skills — pre-LLM RL trained from random init

cephaloform · x · 2026-09-26

The author pushes back on the popular claim that RL can only sharpen capabilities already latent in a model, noting it falls apart when you recall that nearly all pre-LLM-era RL was done on random-init models. He quips that the idea of a random init containing all possible capabilities is "almost poetic."

Original post →

More from Research

Research channel →