Toby Ord Stands by RL Information-Bottleneck Thesis in Exchange with Millidge
Oxford philosopher Toby Ord, in a series of public exchanges with researcher Beren Millidge, reaffirmed his core argument about the information efficiency of reinforcement learning (RL) and offered additional explanations for why RL still works.
Confirmed
- Millidge argued there is an information-theoretic paradox in LLM RL: pretraining computes a loss on every token, whereas policy-gradient RL only receives a scalar reward—often binary—at the end of rollouts spanning hundreds of thousands of tokens. Averaged out, each sample carries very little information (roughly 1/rollout length), yet it still trains LLMs effectively.
- Asked by Millidge whether he had updated his view that "RL doesn't have enough bits," Ord said his position was largely unchanged: rereading his old posts, he felt the argument still holds and many of its predictions have come true.
- Ord reiterated: the upper bound on information learned per unit of FLOP in RL is far lower than in pretraining, which raises two questions—why RL works and how far it can go.
- Ord's explanations include: the available bits, though few, are extremely relevant; part of the answer for why RL has scaled to today's level lies in mid-training; and RL's generalization on non-verifiable tasks has not been as poor as expected.
- Ord also raised a doubt: models may simply be learning what is being tested, so benchmarks may overstate RL's true progress.
Why it matters
- The debate strikes at a fundamental question about the current RL scaling path: if RL's information efficiency is indeed far below pretraining's, there may be a ceiling on how much sustained capability gains it can support.
- If the hypothesis that "benchmarks may overstate RL progress" holds, it would affect how the true capabilities of frontier models are assessed—worth watching for follow-up validation.
2026-09-23 ~ 2026-09-24 · 5 related posts
Primary sources
- Toby Ord on RL's bits problem: 'I haven't changed my mind much' — his jagged-capabilities predictions held up — tobyordoxford ·
- Toby Ord: RL's Low Bits-per-FLOP May Mean Benchmarks Misrepresent Progress — tobyordoxford ·
- How Can LLM RL Work Despite Getting Only 1 Bit per Rollout? A New Essay Tackles the Puzzle — camhowe1729 ·
- [source] How Can LLM RL Work Despite Getting Only 1 Bit per Rollout? A New Essay Tackles the Puzzle — camhowe1729 · 2026-09-23
- [source] Toby Ord on RL's bits problem: 'I haven't changed my mind much' — his jagged-capabilities predictions held up — tobyordoxford · 2026-09-24
- Toby Ord stands by his RL thesis: lower bits-per-FLOP ceiling explains today's jagged AI capabilities — tobyordoxford · 2026-09-24
- Toby Ord: RL's scaling surprise may hinge on mid-training, benchmark gains may overstate progress — tobyordoxford · 2026-09-24
1 near-duplicate retellings: tobyordoxford