SwiLA paper at COLM 2026 switches among multiple linear maps to beat softmax-vs-linear tradeoff

AccBalanced · x · 2026-10-06

In a COLM 2026 paper, the authors propose Switching Linear Attention (SwiLA). The motivation: softmax attention is expressive but its KV cache grows with sequence length, while linear attention keeps a fixed state size at the cost of expressivity. SwiLA aims to get the best of both worlds by dynamically switching among multiple linear maps, keeping the state bounded while recovering expressivity. Details are in the linked thread.

Original post →

More from Research

Research channel →