MIT's Triadic Linear Attention Expands RNN State to 3D Tensors for Better Long-Context Recall
Massachusetts-Institute-of-Technology · hf · 2026-10-05
MIT researchers propose Triadic Linear Attention, generalizing linear attention's state-writing mechanism.
- Instead of a key-value outer product into a matrix state, it writes the triadic outer product of two keys and a value into a third-order (3D) tensor state, reading via contraction of both key axes with two queries.
- An E-dimensional second key yields an E-fold increase in state size while adding only two projections.
- It is compatible with data-dependent forgetting, the delta rule, and chunkwise-parallel training; applied to Gated DeltaNet and scalar-gated linear attention, it substantially improves long-context language modeling and recall, outperforming alternative state-enlargement approaches.
Related event: Triadic Linear Attention Extends Linear RNNs with 3D Tensor States(3 posts)→
More from Research
- Auditing multi-agent collusion talk slated for COLM 2025 workshop — nandofioretto · 2026-10-05
- Fine-tuned 350M model lifts PII removal from 89.7% to 99.7% — JosephJacks_ · 2026-10-05
- Wondersearch Claims It Beat Every Dense Embedding Model on SciFact — Researchers Skeptical — NirantK · 2026-10-05
- Mathematicians plus Meta's Muse Spark solve 6 open math problems, skeptics unmoved — YiMaTweets · 2026-10-05
- Comprehensive Triton GPU programming lecture: from H100 internals to FlashAttention — kalyan_kpl · 2026-10-05
- MICCAI 2026 Accepts 1,167 of 4,402 Papers; FAU Erlangen Lands 11 Plus Two Challenge Wins — maier_ak · 2026-10-05