Smart offloading of active requests erases linear attention's memory advantage
samsja19 · x · 2026-09-09
samsja19 argues a new KV cache offloading scheme differs fundamentally from CPU offloading like Mooncake: Mooncake offloads inactive requests (an artifact of agentic rollouts), while this approach offloads active requests. Since memory usage was the main argument for linear attention over softmax attention, smart offloading erases that advantage.
More from Infra
- Proteus Generates Custom GPU Kernels for Qwen3 122B, Up to 5.2x Faster Than vLLM — matei_zaharia · 2026-09-09
- Google Cloud August AI Infra Roundup: gVisor Sandboxes on Ray, Filestore on Colossus — dl_weekly · 2026-09-09
- Provably private inference services promise prompts unreadable to providers — corbtt · 2026-09-09
- 300B output tokens, $20-30M in compute: Ethan Mollick says AI science will need far more compute — eldonredwards · 2026-09-09
- Rumor resurfaces: Google may replace Nvidia as TSMC's biggest customer — zephyr_z9 · 2026-09-09
- Tahuna open-sources ephemeral GPU orchestration for ML workloads — Monaim101 · 2026-09-09