Model Scaling Trend: 2T Parameters Becoming New Norm as KV Cache Shrinks 10x YoY
zephyr_z9 · x · 2026-09-02
A discussion on AI model scaling trends. The author refutes the idea that parameters are decreasing, arguing the trend is always optimization + more parameters. After a dip post-GPT-4, new models like Fable/Spud, Astra, Kimi, and Qwen 3.8 Max have crossed the 2T parameter threshold. The core debate is whether future models need 100T parameters or if 25T looped transformers suffice. Additionally, while KV cache per token shrinks by at least 10x year-over-year (notably for DeepSeek), context length is steadily increasing from 32k to 1M+.
More from AGI Musings
- AI triggers identity crisis for software engineers as coding roles shift — lee_stott · 2026-09-02
- Reddit pitches "Personal Agent Network": an aligned AI agent negotiating for every person — Shift_Impossible · 2026-09-02
- Users speak 2.5x more to AI agents than humans — ayushtweetshere · 2026-09-02
- Generation is cheap, but physical verification remains the bottleneck in science — shyamalanadkat · 2026-09-02
- Ramez: AI Security Must Secure the Ecosystem, Not Just the Model — alvelda · 2026-09-02
- Joseph Jacks urges more long-term planning for 2040s and 2050s — JosephJacks_ · 2026-09-02