TPA: Tensor Product Attention unifies MHA/GQA/MLA, NeurIPS 2025 Spotlight

burny_tech · x · 2026-09-18

The thread highlights the arXiv paper "Tensor Product Attention Is All You Need" (NeurIPS 2025 Spotlight). Key points:

The quote tweet notes MHA, GQA and MLA can all be seen as special cases of a broader tensor-factorization attention (Tucker Attention), fully compatible with FlashAttention.

Original post →

More from Research

Research channel →