Discussions Heat Up Over KDA Linear Architecture

teortaxesTex · x · 2026-07-18

The author argues that KDA is a "good design," pushing back against the surprise surrounding it: the Kimi team turning it into a product-level result inherently proves extensive research. Furthermore, the 48B version already performed well, making the KDA route entirely unsurprising.

The embedded images focus on the Kimi Linear report and subsequent replication experiments. Some compared KDA against full attention and sliding window attention, finding it performed strongly on tasks like MOAR. Another image summarized the engineering hurdles of KDA compared to full attention, such as implementation, hyperparameter tuning, parallelism, and aligning training/inference numerics. The overarching takeaway is that this linear/alternative attention scheme isn't a "half-baked new architecture," but a robust solution standing on rigorous experimentation and engineering refinement.

Original post →

More from Models

Models channel →