Qwen, Kimi and GLM dropped full attention — 8 attention designs explained
julsimon · x · 2026-10-11
AWS Developer Advocate Julien Simon released a new video on a quiet shift in open-source LLMs.
- Alibaba's Qwen, Moonshot's Kimi and Z.ai's GLM have replaced full attention with cheaper variants in most or all layers — verifiable in their config files.
- The catch: almost nobody has tested these designs against full attention at scale on long tasks. Savings are measured; quality evidence is thin.
- The video covers eight attention designs, who actually tested them, and six checks before building on one. Slides and animations are CC BY 4.0.
More from Models
- RecSys veterans on repeated data overfitting: 'that's why we only train one epoch' — thomasahle · 2026-10-11
- LLMs Reward Information Gain — Just Like the Best Humans Do — sanderssays · 2026-10-11
- Microsoft's decision model promised 80ms, serves 300ms via OpenRouter — DotaMate · 2026-10-11
- Meta's Muse growth slowing: daily active user gains down 62.4% from September surge — AccBalanced · 2026-10-11
- Researcher claims OpenAI exploits user prompts, cites Tao atop a list — basedjensen · 2026-10-11
- 'Nothing new': researcher says engram overfitting is plain overfitting tied to over-parametrization — teortaxesTex · 2026-10-11