How Can LLM RL Work Despite Getting Only 1 Bit per Rollout? A New Essay Tackles the Puzzle
camhowe1729 · x · 2026-09-23
Beren Millidge published a speculative essay on the information-theoretic puzzle of LLM RL: pretraining computes a loss on every token, while policy-gradient RL distills an entire rollout — often hundreds of thousands of tokens, decoded at memory-bandwidth-limited inference cost — into a single, often binary, reward. That works out to roughly 1/rolloutlength bits per sample, and crude reward functions frequently fail rollouts over trivial formatting errors.
By these arguments RL should be doubly inefficient and merely elicit behaviors already present in the base model. Yet empirically the opposite seems true, and the essay explores why RL works anyway, responding to Toby Ord's "RL doesn't have enough bits" argument. The post shares the essay and asks Ord whether he's updated his view.
More from AGI Musings
- 10% of UK Parliament Speeches Are Now AI-Drafted, The Economist Reports — steverathje2 · 2026-09-24
- Yoshua Bengio Addresses UN Security Council on Threat of Uncontrolled Frontier AI Agents — AndrewCritchPhD · 2026-09-24
- OpenAI's Boaz Barak: no benevolent AI dictators, even if it means less abundance — tszzl · 2026-09-24
- CNAS report maps the spectrum of AGI futures and how to shape them — paul_scharre · 2026-09-24
- Philosopher Carissa Veliz: AI dominance is not inevitable — the future is produced, not predicted — CarissaVeliz · 2026-09-24
- Karpathy's 11-month tone shift: from coining vibe coding to feeling far behind — IgorCarron · 2026-09-24