Analysis of Reasoning Scaling Curves: RLVR Compute Efficiency and Generational Leaps

Toby Ord and Ramez Naam analyze new reasoning scaling curves: RLVR yields roughly 3x token efficiency per 10x compute, and new models leap ahead by the equivalent of 100x more tokens. The halved slope may reflect harder datasets, while generational jumps likely stem from math-heavy RL.

2026-09-14 ~ 2026-09-14 · 5 related posts