Toby Ord: RL's scaling surprise may hinge on mid-training, benchmark gains may overstate progress

tobyordoxford · x · 2026-09-24

Toby Ord discusses RL scaling: generalization to non-verifiable tasks is better than expected, possibly due to mid-training; or models may just be learning what's tested, making benchmarks misrepresent overall progress.

Related event: Toby Ord Stands by RL Information-Bottleneck Thesis in Exchange with Millidge(5 posts)→

Original post →

More from Models

Models channel →