ConvRot Quant joins llama-cpp: Q6 accuracy nears Q8 quality

giveen · reddit · 2026-08-24

The ConvRot Quant method is now integrated into llama-cpp-turboquant. Benchmarks show Q6CR achieves KLD/PPL metrics close to Q8, with Q5CR also showing slight improvements. Additionally, a new --moe-cache auto feature helps optimize running MoE models larger than VRAM. Previous decoding and crashing issues noted in PRs have been resolved.

Original post →

More from Infra

Infra channel →