The Extender: log-structured Transformer cuts attention memory 104x

UIChicago · hf · 2026-10-06

The Extender is a log-structured Transformer variant. It adds a concatenation channel x alongside the residual stream: each layer emits a residual update δ and a much smaller extension ε appended to x, and attention KV projections take only x. This shrinks persistent attention memory from 2L·dmodel to the sum of |εell|. With |ε|=32, it matches Transformer accuracy on CORE short-context tasks at 199M–924M params and exceeds it on RULER long-context at 924M; persistent attention memory is 104x smaller than MHA, with savings growing with model width.

Original post →

More from Research

Research channel →