Triadic Linear Attention: extending matrix-state RNNs to a 3D tensor state
ChengleiSi · x · 2026-10-05
Starting from the observation that linear attention is arguably the most naive RNN yet massively outperforms traditional RNNs by maintaining a matrix state, the author proposes a generalization: use a (triadic) outer product of three vectors to maintain a three-dimensional tensor state, dubbed Triadic Linear Attention. The thread expands on the technical details of this new state-space design for linear attention models.
More from Research
- Speech AI leader Shinji Watanabe's h-index hits 100, 26 years after first paper — olexandr · 2026-10-05
- Formalizing long PDE and probability papers now takes just 24-48 hours, says mathematician open-sourcing Lean skills — kfountou · 2026-10-05
- Chalmers: at least 50% credence that augmented LLMs could be conscious within a decade — pickover · 2026-10-05
- RLHF explained in three steps: SFT, reward model, then PPO with a KL penalty — glenbeer · 2026-10-05
- David Duvenaud: giving LLMs the right tools may unlock true reasoning — ThoreG · 2026-10-05
- Blog explores the "shape" of language models and their future tradeoffs in harness design — layer07_yuxi · 2026-10-05