Discussions Heat Up Over KDA Linear Architecture
teortaxesTex · x · 2026-07-18
The author argues that KDA is a "good design," pushing back against the surprise surrounding it: the Kimi team turning it into a product-level result inherently proves extensive research. Furthermore, the 48B version already performed well, making the KDA route entirely unsurprising.
The embedded images focus on the Kimi Linear report and subsequent replication experiments. Some compared KDA against full attention and sliding window attention, finding it performed strongly on tasks like MOAR. Another image summarized the engineering hurdles of KDA compared to full attention, such as implementation, hyperparameter tuning, parallelism, and aligning training/inference numerics. The overarching takeaway is that this linear/alternative attention scheme isn't a "half-baked new architecture," but a robust solution standing on rigorous experimentation and engineering refinement.
More from Models
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11
- Terminal Bench v4: GLM-5.3 Leads at 41.9%, Kimi-K3 Underwhelms at 12.6% — Ok_Warning2146 · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11