Baseten engineer says Kimi K3’s leap came from a chain of targeted fixes, not scale alone
khademinori · x · 2026-07-28
A Baseten engineer spent 48 hours dissecting Kimi K3’s architecture and argued the model’s gains are not just about scale.
- The thread says K3 can “fit” 22,580 GPT-2-sized models from 2019, highlighting how far model scale has advanced in seven years.
- But the core claim is that the progress came from a sequence of targeted fixes, not brute-force parameter growth.
- It traces the path from GPT-2 to linear attention, DeltaNet, gated DeltaNet, and Kimi’s own KDA attention.
- Each iteration is framed as solving a specific weakness of the previous generation: memory limits, information interference, inability to selectively forget, and residual stream dilution.
- The post also argues open source accelerates this kind of architectural scrutiny, because once code is public, outside engineers can quickly map out the technical lineage.
Related event: Deep dives trace Kimi K3 to seven years of LLM architecture evolution(5 posts)→
More from Models
- OpenAI’s Codex now splits into Sol, Terra, and Luna, with Luna priced at one-fifth of Sol — TinfoilTricorn · 2026-07-28
- Alibaba launches the Qwen3.8 Growth Plan after developer feedback on Qwen3.8-Max-Preview — Alibaba_Qwen · 2026-07-28
- Kimi K3 license keeps MIT terms but adds revenue and user-count restrictions — TheZachMueller · 2026-07-28
- OpenRouter data covers under 1% of global inference, so it can’t prove Chinese models lead usage — zephyr_z9 · 2026-07-28
- GLM-5.2 runs locally on Dell Pro Max at 40 tokens/s, hinting at a new distillation pipeline — pcuenq · 2026-07-28
- DesignArena leak points to two new GPT codenames, Zinc and Magnesium — TheZachMueller · 2026-07-28