Kimi team says delta-rule state space models are more expressive than attention

teortaxesTex · x · 2026-07-26

Kimi team argues that state space models with the delta rule are strictly more expressive than attention, challenging the idea that attention is all you need.

The post also suggests this may be easier to train end to end than the current cascade of sparsity tricks used in V4, and hopes both approaches can work well.

Original post →

More from Research

Research channel →