Model Scaling Trend: 2T Parameters Becoming New Norm as KV Cache Shrinks 10x YoY

zephyr_z9 · x · 2026-09-02

A discussion on AI model scaling trends. The author refutes the idea that parameters are decreasing, arguing the trend is always optimization + more parameters. After a dip post-GPT-4, new models like Fable/Spud, Astra, Kimi, and Qwen 3.8 Max have crossed the 2T parameter threshold. The core debate is whether future models need 100T parameters or if 25T looped transformers suffice. Additionally, while KV cache per token shrinks by at least 10x year-over-year (notably for DeepSeek), context length is steadily increasing from 32k to 1M+.

Original post →

More from AGI Musings

AGI Musings channel →