Toby Ord: RL's Low Bits-per-FLOP May Mean Benchmarks Misrepresent Progress

tobyordoxford · x · 2026-09-24

Reacting to Beren Millidge's new essay, Toby Ord restated his core argument: RL has a much lower ceiling on information learnable per FLOP than pretraining, raising questions about how it works and how far it can go. He adds possible resolutions — the bits may be highly relevant, and RL may only be learning what is tested, meaning benchmarks could be misrepresenting overall progress. He welcomed Millidge's deeper explanations.

Related event: Toby Ord Stands by RL Information-Bottleneck Thesis in Exchange with Millidge(5 posts)→

Original post →

More from AGI Musings

AGI Musings channel →