SwiLA paper at COLM 2026 switches among multiple linear maps to beat softmax-vs-linear tradeoff
AccBalanced · x · 2026-10-06
In a COLM 2026 paper, the authors propose Switching Linear Attention (SwiLA). The motivation: softmax attention is expressive but its KV cache grows with sequence length, while linear attention keeps a fixed state size at the cost of expressivity. SwiLA aims to get the best of both worlds by dynamically switching among multiple linear maps, keeping the state bounded while recovering expressivity. Details are in the linked thread.
More from Research
- Survey: 65% of Japanese seniors prefer robot-assisted nursing homes, willing to pay 8% more — HealthcareLdr · 2026-10-06
- AI Scholar Yi Ma Proposes Rolling 10-Paper Cap on arXiv to Curb Paper Flooding — YiMaTweets · 2026-10-06
- FlashDexRetarget: one RL policy retargets hand demos at 90% success, ~100x less compute — KyleMorgenstein · 2026-10-06
- Stanford ACE team unveils Sentry: failure tips in context hurt LLM agents, +39% gains — StanfordAILab · 2026-10-06
- From Kaggle Champion to ULMFiT: How Jeremy Howard Rewrote Language Model Training — bigaiguy · 2026-10-06
- Training on a Post-Trained Model Often 'Fries' It, Causing Reality Drift — Sauers_ · 2026-10-06