Asking for real quality benchmarks of Beellama's old-cache-only quantization for local coding LLMs

RadianceTower · reddit · 2026-09-13

A Reddit user asks for proper quality benchmarks of Beellama (and its fork Beellama-kvarn), a local inference technique that quantizes only older KV cache while keeping recent context at higher precision — potentially better than uniformly quantizing all cache. The recommended 1k tail raises questions (why not 20k on a 240k context?). Beellama-kvarn claims further performance gains pending merging upstream. The open question: how much does this hurt coding quality? No authoritative benchmarks yet.

Original post →

More from Infra

Infra channel →