How Can LLM RL Work Despite Getting Only 1 Bit per Rollout? A New Essay Tackles the Puzzle

camhowe1729 · x · 2026-09-23

Beren Millidge published a speculative essay on the information-theoretic puzzle of LLM RL: pretraining computes a loss on every token, while policy-gradient RL distills an entire rollout — often hundreds of thousands of tokens, decoded at memory-bandwidth-limited inference cost — into a single, often binary, reward. That works out to roughly 1/rolloutlength bits per sample, and crude reward functions frequently fail rollouts over trivial formatting errors.

By these arguments RL should be doubly inefficient and merely elicit behaviors already present in the base model. Yet empirically the opposite seems true, and the essay explores why RL works anyway, responding to Toby Ord's "RL doesn't have enough bits" argument. The post shares the essay and asks Ord whether he's updated his view.

Related event: Toby Ord Stands by RL Information-Bottleneck Thesis in Exchange with Millidge(5 posts)→

Original post →

More from AGI Musings

AGI Musings channel →