Toby Ord: 10x more RLVR compute cuts tokens-to-target ~3x; gains may be math-specific
tobyordoxford · x · 2026-09-14
Oxford philosopher Toby Ord replies to Ramez Naam with a rule of thumb he'd noticed: every 10x increase in RLVR compute yields roughly a 3x reduction in the tokens needed to reach the same standard (with paper link).
Naam responds that the dramatic jump between model generations pays for 2-3 orders of magnitude of tokens, and guesses the new model was heavily RL'd on formal math — meaning the gains may be specific to mathematics. Both remain cautious about whether the capability leap generalizes beyond math; "we'll find out eventually."
Related event: Oxford's Toby Ord: 10x Compute in RLVR Yields ~3x Token Efficiency(3 posts)→
More from Models
- David Bellamy clarifies his experiment used K2 Horizon, an open-weights 375B LLM — JeremyNguyenPhD · 2026-09-14
- Swift-Qwen3.8-27b, a token-efficient reasoning Qwen finetune, trends on Hugging Face — ukisai · 2026-09-14
- GPT-6 Astra hands-on: composes first, orchestrates later, and reportedly outshines Fable and Sol — paw_lean · 2026-09-14
- Researcher posts proof he both synthesized viruses and trained a 375B open-weight LLM — ethanCaballero · 2026-09-14
- New scaling curve has half the slope: 10,000x compute for 20%-to-80%, but bigger generational jumps — tobyordoxford · 2026-09-14
- Fudan NLP paper explains why max reasoning settings can backfire on SWE benchmarks — karminski3 · 2026-09-14