New T² Scaling Law Says Chinchilla's 20 Tokens/Param Is Wrong in the Test-Time Inference Era
josh_wills · x · 2026-09-05
Nicholas Roberts, in a talk hosted by Datology AI, introduces the Train-to-Test (T²) law, which folds pass@k into pretraining scaling: once models reason at test time, the compute-optimal strategy shifts toward small models trained on far more data.
- A 350M model overtrained to 80,000+ tokens per parameter looks wasteful but is optimal under T² scaling
- Chinchilla's 20 tokens/param is the wrong target in the test-time reasoning era
- The catch: you hit the data wall much sooner, making data quality far more important
More from Research
- Extreme reward functions push LLMs away from their current behavior — jessi_cata · 2026-09-05
- VeriPhy: agentic physical reasoning framework for world model evaluation — Wenzhuo Xu · 2026-09-05
- TRACES agent benchmark grades live execution loops, not answers — SucceededMind · 2026-09-05
- Pedro Domingos quips: 'new idea' called RNNs will power next-gen LLMs — pmddomingos · 2026-09-05
- LoRA Creator Edward Hu Publishes Guide on Post-Training Open-Source Models with RL — iamrobotbear · 2026-09-05
- Amid CoT monitoring buzz, one video offers a glimpse into how LLMs actually think — kastnerkyle · 2026-09-05