Kimi team says delta-rule state space models are more expressive than attention
teortaxesTex · x · 2026-07-26
Kimi team argues that state space models with the delta rule are strictly more expressive than attention, challenging the idea that attention is all you need.
The post also suggests this may be easier to train end to end than the current cascade of sparsity tricks used in V4, and hopes both approaches can work well.
More from Research
- Qwen3-8B fine-tune lost its thinking mode because the template taught empty `<think>` blocks — MeldhLLC · 2026-07-26
- Enterprises should own business knowledge, not models, as frontier agents double every 6–7 months — hamostaf04 · 2026-07-26
- GEO course roundup adds datasets, tools and white papers for AI citation optimization — vista8 · 2026-07-26
- Governance graphs cut multi-agent collusion from 50% to 5.6% in a new study — sebkrier · 2026-07-26
- Researchers debate how to do science after capable LLMs cross a threshold — MvsCerezo · 2026-07-26
- AI helps solve a major quantum information problem on Werner state distillability — MvsCerezo · 2026-07-26