Smart offloading of active requests erases linear attention's memory advantage

samsja19 · x · 2026-09-09

samsja19 argues a new KV cache offloading scheme differs fundamentally from CPU offloading like Mooncake: Mooncake offloads inactive requests (an artifact of agentic rollouts), while this approach offloads active requests. Since memory usage was the main argument for linear attention over softmax attention, smart offloading erases that advantage.

Original post →

More from Infra

Infra channel →