Kimi K3's KDA vs Qwen3.8's GDN: channel-wise gating enables 2D rotations GDN can't
orvieto_antonio · x · 2026-09-23
Delta-rule linear attention goes mainstream
- Delta-rule linear attention now ships in production LLMs: Qwen3.8 uses Gated DeltaNet (GDN), while Kimi K3 uses Kimi Delta Attention (KDA).
- Key difference: KDA replaces GDN's scalar gate with a channel-wise gate; the author asks whether this actually improves recurrence expressivity.
- Finding: with just two tweaks, KDA can perform 2D rotations — something GDN cannot do — showing the channel-wise gate brings genuinely richer structure.
More from Research
- Nature Health paper: AI is now a determinant of health — time for an epidemiology of AI — EricTopol · 2026-09-23
- ProgramBench: rebuilding programs from binaries is brutal — Claude Opus 5 leads at 4.5% resolved — jyangballin · 2026-09-23
- Building AI agent harnesses that improve across 6 editable control dimensions — MaryamMiradi · 2026-09-23
- DeepMind Tests Whether RLAIF Can Match Human Feedback Across Summarization and Dialogue Tasks — burkov · 2026-09-23
- Braidwell Founders in TIME: AI Accelerating Science, Promise Lies in People — AndrewLBeam · 2026-09-23
- YC-backed ORO launches ORO Bench, a synthetic shopping benchmark powering Bittensor SN15 — markjeffrey · 2026-09-23