Model quant vs KV quant: Which tradeoff is better?
Elorun · reddit · 2026-08-20
A user is weighing two quantization setups for Qwen 3.8 27B on llama.cpp while maintaining a context length above 150k: Q4KM weights with Q8 KV cache, or Q4KXL weights with Q4 KV cache. The question is whether it's better to use larger weight quantization with lower KV cache precision, or smaller weight quantization with higher KV cache precision.
More from Infra
- Llama-Mobile: 2.7-Bit Quantization Shrinks Llama 3.2 Vision 11B to 3.7GB for Phones — Luka Ribar · 2026-08-24
- DSCO Router Launches Unified Gateway for Multi-Model Routing with BYOK Support — arthurcolle · 2026-08-24
- Open Source RobotSoul: Persistent Identity for Agents After Context Resets — robauto-dot-ai · 2026-08-24
- Offloading MoE models to RAM causes slow prefill speeds — former_farmer · 2026-08-24
- Etched Raises $1B Led by Jane Street to Validate Architecture-Agnostic AI Chips — TheTuringPost · 2026-08-24
- ConvRot Quant joins llama-cpp: Q6 accuracy nears Q8 quality — giveen · 2026-08-24