Block KV cache streaming PR bounds VRAM at long context in llama.cpp fork
giveen · reddit · 2026-09-06
Developer giveen submitted PR #357 to llama-cpp-turboquant implementing block KV cache streaming via a shared CUDA phase arena, bounding VRAM usage at long context. The author credits Raymond Huang's llama.cpp-adaptive-kv-streaming as the original work, ported it to turboX, extended model support beyond Qwen, and benchmarked thoroughly to confirm the gains.
More from Infra
- T-Glass shortage worsens: Kinsus losing 10-15% of monthly ABF revenue, 25% capacity expansion planned for 2027 — zephyr_z9 · 2026-09-06
- Bump-less 3D stacking goes practical: Intel Diamond Rapids first, AMD Zen rumored next — bookwormengr · 2026-09-06
- Self-hosting AI: what rigs do local LLM runners actually use? — Fun_Kangaroo512 · 2026-09-06
- 200 tok/s on 8GB VRAM: dev benchmarks 6 small models for local AI — TheMoonMidas · 2026-09-06
- Reddit debate: is compute the real hurdle to automating all cognitive labour? — Vivid-Flamingo-644 · 2026-09-06
- Huawei Chief Scientist Details Logic Folding Cooling: Density Near TSMC N3E Level — teortaxesTex · 2026-09-06