Can llama.cpp share KV cache across multiple GPUs for parallel requests?
spaceman_ · reddit · 2026-08-22
User explores running Qwen3.8-27B on 4x AMD R9700 GPUs. Tests show PCIe bottlenecks prevent speedup for single-request splitting. The challenge is that multiple llama-server instances cannot share KV cache; routing a session's second request to a different GPU invalidates the cache. User asks if there is a setting to load weights across cards while sharing a unified KV cache.
More from Infra
- Ox Alpha's 100T tokens/day giveaway costs $1-2M in electricity at full capacity — teortaxesTex · 2026-08-22
- Seeking unified AI gateway for OpenAI cost visibility — HurryOrganic · 2026-08-22
- Agents fail silently: $47k loop reveals monitoring gaps — alifgokce · 2026-08-22
- Anti-datacenter movement is humanity recognizing its successor — ZeroStateReflex · 2026-08-22
- FreeToken: 4x Faster Decode, Enables 284B Models on Gaming Desktops — airesearch12 · 2026-08-22
- GLM-5.2 local inference: ubatch size significantly boosts MoE performance — fuzhongkai · 2026-08-22