New generation's jump equals 100x tokens on old model, like four generations in the o1-to-GPT-5 era

tobyordoxford · x · 2026-09-14

In a thread with ramez on inference-time scaling, tobyordoxford notes the new generation's jump equals what the old model would achieve with 100x more tokens. The log-scaling slope itself hasn't improved, but the between-generation jump is huge — comparable to four generations in the o1→o3→GPT-5 era. He also cites his rule of thumb: each 10x more RLVR data cut required tokens by 3x.

Related event: Toby Ord Analyzes New Reasoning Scaling Curves: Halved Slope but Massive Generational Leaps(5 posts)→

Original post →

More from Models

Models channel →