Memory Attention: Paper Replaces Value Projection With Lookup, Cuts GPU Storage 7.38%

serrjoa · x · 2026-09-27

A new alphaXiv paper by Jiale Kang proposes Memory Attention (MA), replacing the attention value projection with a token-indexed lookup.

Core idea

Engineering highlights

Under matched training token budgets, experiments across attention configurations show improved language modeling and downstream performance. The author, inspired by Engram, frames this as a candidate direction for future model architectures.

Original post →

More from Infra

Infra channel →