Toby Ord: RL's Low Bits-per-FLOP May Mean Benchmarks Misrepresent Progress
tobyordoxford · x · 2026-09-24
Reacting to Beren Millidge's new essay, Toby Ord restated his core argument: RL has a much lower ceiling on information learnable per FLOP than pretraining, raising questions about how it works and how far it can go. He adds possible resolutions — the bits may be highly relevant, and RL may only be learning what is tested, meaning benchmarks could be misrepresenting overall progress. He welcomed Millidge's deeper explanations.
More from AGI Musings
- 10% of UK Parliament Speeches Are Now AI-Drafted, The Economist Reports — steverathje2 · 2026-09-24
- Yoshua Bengio Addresses UN Security Council on Threat of Uncontrolled Frontier AI Agents — AndrewCritchPhD · 2026-09-24
- OpenAI's Boaz Barak: no benevolent AI dictators, even if it means less abundance — tszzl · 2026-09-24
- CNAS report maps the spectrum of AGI futures and how to shape them — paul_scharre · 2026-09-24
- Philosopher Carissa Veliz: AI dominance is not inevitable — the future is produced, not predicted — CarissaVeliz · 2026-09-24
- Karpathy's 11-month tone shift: from coining vibe coding to feeling far behind — IgorCarron · 2026-09-24