A long history of LLM architectures leads from GPT-2 to Kimi K3
baseten · x · 2026-07-28
A long-form writeup traces the history of LLM architectures from 2017 to Kimi K3.
- The post frames Kimi K3 as the endpoint of a long scaling and architecture evolution rather than a standalone release.
- In the quoted article, the author says 22,580 GPT-2-sized models would fit inside one Kimi K3-sized system, using that contrast to emphasize how far model scale has moved since 2019.
- The thread is essentially a technical reading recommendation for people who want a broader architectural timeline, not just a product announcement.
More from Research
- Nature paper measures non-Gaussian order-parameter statistics across a phase transition — burny_tech · 2026-07-29
- Kimi K3 Tech Report: How Moonshot Achieved 2.5x Compute Efficiency — alex_verem · 2026-07-29
- RG view of generalization says neural nets learn scale-invariant correlation structure — burny_tech · 2026-07-29
- Quanta profiles 2026 Fields Medalist Yu Deng and his meticulous research style — burny_tech · 2026-07-29
- New scaling law paper says repetition can beat paraphrasing for some pretraining regimes — burny_tech · 2026-07-29
- Replication finds agent experience distillation preserves 44.1% of ICL gains on SWE tasks — burny_tech · 2026-07-29