ConvRot Quant joins llama-cpp: Q6 accuracy nears Q8 quality
giveen · reddit · 2026-08-24
The ConvRot Quant method is now integrated into llama-cpp-turboquant. Benchmarks show Q6CR achieves KLD/PPL metrics close to Q8, with Q5CR also showing slight improvements. Additionally, a new --moe-cache auto feature helps optimize running MoE models larger than VRAM. Previous decoding and crashing issues noted in PRs have been resolved.
More from Infra
- Intel Expects Wins with MSFT, QCOM, MRVL on ASIC, CPU, CPO — BenBajarin · 2026-08-24
- Llama-Mobile: 2.7-Bit Quantization Shrinks Llama 3.2 Vision 11B to 3.7GB for Phones — Luka Ribar · 2026-08-24
- DSCO Router Launches Unified Gateway for Multi-Model Routing with BYOK Support — arthurcolle · 2026-08-24
- Open Source RobotSoul: Persistent Identity for Agents After Context Resets — robauto-dot-ai · 2026-08-24
- Offloading MoE models to RAM causes slow prefill speeds — former_farmer · 2026-08-24
- Etched Raises $1B Led by Jane Street to Validate Architecture-Agnostic AI Chips — TheTuringPost · 2026-08-24