Linear attention takes up to 75% of decode latency at large batch, authors say

sehoonkim418 · x · 2026-10-08

The authors give context: in hybrid LLMs, linear attention reads the full recurrent state from GPU memory each step, taking up to 75% of decode latency at large batch. ReplaySSM buffers updates and writes the state once per multi-step window, but reads remain — linear attention still takes up to 53%.

Related event: Linear Attention State Reads Consume Up to 75% of Large-Batch Decode Latency(2 posts)→

Original post →

More from Infra

Infra channel →