Linear attention takes up to 75% of decode latency at large batch, authors say
sehoonkim418 · x · 2026-10-08
The authors give context: in hybrid LLMs, linear attention reads the full recurrent state from GPU memory each step, taking up to 75% of decode latency at large batch. ReplaySSM buffers updates and writes the state once per multi-step window, but reads remain — linear attention still takes up to 53%.
More from Infra
- Only Samsung HBM meets Nvidia Vera Rubin performance requirements, per leak — zephyr_z9 · 2026-10-08
- Box CEO on agent compute: one app serving 100M users would need $2.8B in infra — inductionheads · 2026-10-08
- Firmus, valued near $44bn, may shelve ASX IPO as investors balk — nordicinst · 2026-10-08
- DWDM wavelength lock tightens from ±12.5GHz to ±3.5GHz, pushing optical interconnect costs — jwt0625 · 2026-10-08
- llama.cpp merges Metal kernel PR covering all 26 weight formats, up to 4.4x faster MMA on Mac — ggerganov · 2026-10-08
- Theo: viral $132M/year token cost claim is wrong — closer to $3M now, $1.2k soon — dotey · 2026-10-08