Analysis of Reasoning Scaling Curves: RLVR Compute Efficiency and Generational Leaps
Toby Ord and Ramez Naam analyze new reasoning scaling curves: RLVR yields roughly 3x token efficiency per 10x compute, and new models leap ahead by the equivalent of 100x more tokens. The halved slope may reflect harder datasets, while generational jumps likely stem from math-heavy RL.
2026-09-14 ~ 2026-09-14 · 5 related posts
- New scaling curve has half the slope: 10,000x compute for 20%-to-80%, but bigger generational jumps — tobyordoxford · 2026-09-14
- ramez: generational jumps pay for 2-3 orders of magnitude of tokens, suspects heavy math RL — ramez · 2026-09-14
- Toby Ord: 10x more RLVR compute cuts tokens-to-target ~3x; gains may be math-specific — tobyordoxford · 2026-09-14
- New generation's jump equals 100x tokens on old model, like four generations in the o1-to-GPT-5 era — tobyordoxford · 2026-09-14
- Halved slope may be an artifact of higher-difficulty-variance dataset, log scaling still holds — tobyordoxford · 2026-09-14