Kimi K3 paper drops positional embeddings and departs from Transformer orthodoxy
bookwormengr · x · 2026-07-28
The author says Kimi K3’s paper contains a surprising number of real innovations and departures from transformer orthodoxy, arguing that it is not just “mostly a Transformer with a few Noam mods.”
They highlight one especially bold choice: removing positional encodings/embeddings entirely, and say they’ll share more thoughts after a deeper read.
More from Research
- New scaling law paper says repetition can beat paraphrasing for some pretraining regimes — burny_tech · 2026-07-29
- Replication finds agent experience distillation preserves 44.1% of ICL gains on SWE tasks — burny_tech · 2026-07-29
- Cohere Labs brings a Triton tutorial and agentic AI panel to Deep Learning Indaba — Cohere_Labs · 2026-07-29
- UltraEP brings real-time MoE balancing to Xiaohongshu’s rack-scale inference stack — 机器之心 · 2026-07-29
- Triangle Splatting SLAM uses differentiable triangles for dense RGB-D mapping — janusch_patas · 2026-07-29
- DLBCN 2026 opens presenter, spotlight and poster calls for Barcelona deep learning research — serrjoa · 2026-07-29