Kimi K3 paper drops position embeddings and pushes beyond Transformer orthodoxy
peterjliu · x · 2026-07-28
The author says Kimi K3’s paper appears full of genuine architectural departures rather than incremental tweaks.
The standout claim in the post is that the model removes position encodings/embeddings altogether, which the author describes as a bold break from Transformer orthodoxy and evidence that the team is doing real research rather than copying existing designs.
More from Research
- Nature paper measures non-Gaussian order-parameter statistics across a phase transition — burny_tech · 2026-07-29
- Kimi K3 Tech Report: How Moonshot Achieved 2.5x Compute Efficiency — alex_verem · 2026-07-29
- RG view of generalization says neural nets learn scale-invariant correlation structure — burny_tech · 2026-07-29
- Quanta profiles 2026 Fields Medalist Yu Deng and his meticulous research style — burny_tech · 2026-07-29
- New scaling law paper says repetition can beat paraphrasing for some pretraining regimes — burny_tech · 2026-07-29
- Replication finds agent experience distillation preserves 44.1% of ICL gains on SWE tasks — burny_tech · 2026-07-29