ramez: generational jumps pay for 2-3 orders of magnitude of tokens, suspects heavy math RL
ramez · x · 2026-09-14
Responding to tobyordoxford's scaling-curve analysis, ramez agrees the worse slope is interesting but argues the dramatic between-generation jump pays for 2-3 orders of magnitude of tokens. He suspects the model was heavily RL-trained on formal math, making gains possibly math-specific. Key point: bigger baseline jumps can offset declining inference-scaling returns.
More from Models
- Playing GeoGuessr against DeepSeek 4.1 Flash: it's fairly good at geolocation — PMinervini · 2026-09-14
- DeepSeek V4.1 benches low but feels top-tier at 4x speed, insiders say — teortaxesTex · 2026-09-14
- Best small LLMs for writing on 4GB VRAM? Reddit thread weighs the options — Mysterious-Comment94 · 2026-09-14
- Leaked Astra and Fable sizes are far below 10T, says X user: scale barely maps to capability now — teortaxesTex · 2026-09-14
- Free ChatGPT solves a decade-old maths dice problem in 13 minutes — Chris_Armstrong · 2026-09-14
- David Bellamy clarifies his experiment used K2 Horizon, an open-weights 375B LLM — JeremyNguyenPhD · 2026-09-14