SwiLA explained: a mixture of J linear regressions that reduces to DeltaNet at J=1
YouJiacheng · x · 2026-10-06
Researcher hyundongleee breaks down the design of SwiLA, and You Jiacheng highlights that "each value coordinate picks one" makes intuitive sense.
SwiLA's two key choices:
- Function class: a mixture of J linear regressions; each value coordinate picks one, with a learned prior.
- Optimizer: online SGD on a likelihood-based loss; each regressor takes a delta-rule step weighted by its responsibility.
Notably, setting J=1 recovers DeltaNet exactly, giving linear-attention-style methods a unified mixture interpretation.
More from Research
- A regex script scores 95% on User Sim Index, researcher warns eval is broken — ericzelikman · 2026-10-06
- Why User-Model Evals Are Hard: Stanford Researchers Bet on a Multi-User Turing Test — alexisjross · 2026-10-06
- Embedding Every Font with Neural Networks Yields a Flower-Shaped Map of Google Fonts — Chroma-Crash · 2026-10-06
- Crawler Zoo Launches a Free Arena for Testing Local-Model Agents — Time_Instruction_955 · 2026-10-06
- Trained agentic context management: 8K-context small model matches GPT-5.4 at 1M on OOLONG — xennygrimmato_ · 2026-10-06
- User Sim Index is broken: trivial bot scores 95% across behavioral dims — ericzelikman · 2026-10-06