Kimi K3 paper points to hybrid attention, Kimi Linear, and no positional embeddings

peterjliu · x · 2026-07-28

Kimi K3 appears to use hybrid attention and a Kimi Linear variant

The post says the Kimi K3 paper suggests the model uses a hybrid linear / full-attention design and a proprietary Kimi Linear variant, with claims that it outperforms full attention.

The author argues this may be the first confirmed case of a frontier model adopting hybrid attention, and notes how unusual it is to see such departures from standard Transformer orthodoxy in a mainline frontier model. They also point out that the paper appears to remove positional encodings/embeddings.

The implication is that Moonshot is betting on small-scale scaling results strongly enough to ship architectural choices that other teams have struggled to justify.

Related event: Deep Dive into Kimi K3 Architecture: 2.8T Parameters and Attention Innovations(24 posts)→

Original post →

More from Models

Models channel →