The Extender: log-structured Transformer cuts attention memory 104x
UIChicago · hf · 2026-10-06
The Extender is a log-structured Transformer variant. It adds a concatenation channel x alongside the residual stream: each layer emits a residual update δ and a much smaller extension ε appended to x, and attention KV projections take only x. This shrinks persistent attention memory from 2L·dmodel to the sum of |εell|. With |ε|=32, it matches Transformer accuracy on CORE short-context tasks at 199M–924M params and exceeds it on RULER long-context at 924M; persistent attention memory is 104x smaller than MHA, with savings growing with model width.
More from Research
- CC-Bench at COLM 2026 finds LLMs still default to stereotypes over implicit cultural cues — MaartenSap · 2026-10-06
- UCLA PhD's LLM agents produce 126K-line Lean 4 proof of MIP* = RE core theorem in 63 days — siyan_zhao · 2026-10-06
- UniReps 4th edition heads to Paris on Dec 12, 2026, with submissions extended to Oct 10 — ClementineDomi6 · 2026-10-06
- AI-generated key-value stores beat general-purpose DBs via specialized architecture — CShorten30 · 2026-10-06
- Real2sim's last mile for robotics: verification pipeline for physics and affordance — chris_j_paxton · 2026-10-06
- LenVM lands COLM 2026 Spotlight: value model predicts remaining generation length for efficient reasoning — xwang_lk · 2026-10-06