New generation's jump equals 100x tokens on old model, like four generations in the o1-to-GPT-5 era
tobyordoxford · x · 2026-09-14
In a thread with ramez on inference-time scaling, tobyordoxford notes the new generation's jump equals what the old model would achieve with 100x more tokens. The log-scaling slope itself hasn't improved, but the between-generation jump is huge — comparable to four generations in the o1→o3→GPT-5 era. He also cites his rule of thumb: each 10x more RLVR data cut required tokens by 3x.
More from Models
- Playing GeoGuessr against DeepSeek 4.1 Flash: it's fairly good at geolocation — PMinervini · 2026-09-14
- DeepSeek V4.1 benches low but feels top-tier at 4x speed, insiders say — teortaxesTex · 2026-09-14
- Best small LLMs for writing on 4GB VRAM? Reddit thread weighs the options — Mysterious-Comment94 · 2026-09-14
- Leaked Astra and Fable sizes are far below 10T, says X user: scale barely maps to capability now — teortaxesTex · 2026-09-14
- Free ChatGPT solves a decade-old maths dice problem in 13 minutes — Chris_Armstrong · 2026-09-14
- David Bellamy clarifies his experiment used K2 Horizon, an open-weights 375B LLM — JeremyNguyenPhD · 2026-09-14