Kimi team says delta-rule state space models are more expressive than attention
teortaxesTex · x · 2026-07-26
Kimi team argues that state space models with the delta rule are strictly more expressive than attention, challenging the idea that attention is all you need.
The post also suggests this may be easier to train end to end than the current cascade of sparsity tricks used in V4, and hopes both approaches can work well.
Related event: Kimi K3 Architectural Innovations Spark DRAM Demand Debate(5 posts)→
More from Research
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11
- Fruit fly brain as an LLM: connectome-driven language model demo goes live — ngxson · 2026-09-11
- Harry Collins: LLMs can't do frontier science because they can't invent new language — whoamisri · 2026-09-11