Model quant vs KV quant: Which tradeoff is better?

Elorun · reddit · 2026-08-20

A user is weighing two quantization setups for Qwen 3.8 27B on llama.cpp while maintaining a context length above 150k: Q4KM weights with Q8 KV cache, or Q4KXL weights with Q4 KV cache. The question is whether it's better to use larger weight quantization with lower KV cache precision, or smaller weight quantization with higher KV cache precision.

Original post →

More from Infra

Infra channel →