Toby Ord: 10x more RLVR compute cuts tokens-to-target ~3x; gains may be math-specific

tobyordoxford · x · 2026-09-14

Oxford philosopher Toby Ord replies to Ramez Naam with a rule of thumb he'd noticed: every 10x increase in RLVR compute yields roughly a 3x reduction in the tokens needed to reach the same standard (with paper link).

Naam responds that the dramatic jump between model generations pays for 2-3 orders of magnitude of tokens, and guesses the new model was heavily RL'd on formal math — meaning the gains may be specific to mathematics. Both remain cautious about whether the capability leap generalizes beyond math; "we'll find out eventually."

Related event: Oxford's Toby Ord: 10x Compute in RLVR Yields ~3x Token Efficiency(3 posts)→

Original post →

More from Models

Models channel →