Asking for real quality benchmarks of Beellama's old-cache-only quantization for local coding LLMs
RadianceTower · reddit · 2026-09-13
A Reddit user asks for proper quality benchmarks of Beellama (and its fork Beellama-kvarn), a local inference technique that quantizes only older KV cache while keeping recent context at higher precision — potentially better than uniformly quantizing all cache. The recommended 1k tail raises questions (why not 20k on a 240k context?). Beellama-kvarn claims further performance gains pending merging upstream. The open question: how much does this hurt coding quality? No authoritative benchmarks yet.
More from Infra
- MHA, MQA, GQA and MLA explained by what happens to the K/V cache during decoding — techNmak · 2026-09-13
- Early OpenAI employee says 'winning' AGI is outdated — 99% of future compute will run locally — GregCook2011 · 2026-09-13
- Benchmarks show ComfyUI in Docker (CUDA 12.4) runs at 0% penalty if you fix the /dev/shm OOM crash — fluxdraw · 2026-09-13
- Qwen3.8-27B EXL3 one-click kit brings quality local LLM to 16-32GB consumer GPUs — udmrzn · 2026-09-13
- Nex-N2.5-mini-MLX-4bit hits 133.6 tok/s on Apple M5 Max — DerTomsn · 2026-09-13
- Positron shipped an AI chip in 15 months; Ainek rumored at 13 — DavidBennett__ · 2026-09-13