Block KV cache streaming PR bounds VRAM at long context in llama.cpp fork

giveen · reddit · 2026-09-06

Developer giveen submitted PR #357 to llama-cpp-turboquant implementing block KV cache streaming via a shared CUDA phase arena, bounding VRAM usage at long context. The author credits Raymond Huang's llama.cpp-adaptive-kv-streaming as the original work, ported it to turboX, extended model support beyond Qwen, and benchmarked thoroughly to confirm the gains.

Original post →

More from Infra

Infra channel →