Kimi K3 paper points to hybrid attention, Kimi Linear, and no positional embeddings
peterjliu · x · 2026-07-28
Kimi K3 appears to use hybrid attention and a Kimi Linear variant
The post says the Kimi K3 paper suggests the model uses a hybrid linear / full-attention design and a proprietary Kimi Linear variant, with claims that it outperforms full attention.
The author argues this may be the first confirmed case of a frontier model adopting hybrid attention, and notes how unusual it is to see such departures from standard Transformer orthodoxy in a mainline frontier model. They also point out that the paper appears to remove positional encodings/embeddings.
The implication is that Moonshot is betting on small-scale scaling results strongly enough to ship architectural choices that other teams have struggled to justify.
More from Models
- LiquidAI’s 230M LFM2.5 encoder trends on Hugging Face — LiquidAI · 2026-07-29
- Moonshot’s Kimi K3 is a 2.8T open-weight MoE model with 1M-token context — alex_verem · 2026-07-29
- Kimi K3 Tech Report: How Moonshot Achieved 2.5x Compute Efficiency — alex_verem · 2026-07-29
- Experienced web developer says Claude Opus 4.8 and 5 are now unusable for chat — FuzzyHead455 · 2026-07-29
- Reddit user says paid Gemini Pro access is still routing to other models — IAmMonke2 · 2026-07-29
- A fake Claude “model welfare” leak turns into an AI-community meme — repligate · 2026-07-29