BeeLlama.cpp v0.4.1 adds KV-cache precision tails and new quantization modes
Anbeeld · reddit · 2026-07-27
BeeLlama.cpp v0.4.1 adds several KV-cache optimizations for llama.cpp forks, including KVarN, a KV-cache precision tail, and new standard quantization types such as q20/q31 and q60/q61.
The author claims the new settings can significantly reduce VRAM while keeping benchmark quality high. In the published KLD tables for Qwen 3.6 27B Q5KS 64k, some KVarN variants preserve most of the q80 quality at materially lower memory cost, while also documenting the trade-offs for different tails and precision levels.
The post also notes an important caveat: for SWA architectures like Gemma and GPT-OSS, KVarN and KVPT have the same precision but higher VRAM and performance overhead because of ring/mixed-precision complications.
More from Infra
- Triton backend pushes Falcon3-10B to 97.5 tok/s on an RTX 5070 — OCV_Researcher · 2026-07-27
- Sparrow switches its Standard mode to Ministral 3 14B for local document extraction — andrejusb · 2026-07-27
- Nvidia supplier Wistron opens $700 million Texas plant for GB300 and Vera Rubin systems — Beth_Kindig · 2026-07-27
- Running 100 million tokens through GLM 5.2 NVFP4 locally costs about $1 — _akhaliq · 2026-07-27
- Cloudflare’s AI-training block can also stop Googlebot after September 15 — daluoseo · 2026-07-27
- Self-hosted proxy unifies 424 AI models behind one OpenAI-compatible endpoint — ranadheer535 · 2026-07-27