Qwen 3.8 27B tests show KV Cache f16 outperforms q8_0 in quality

Felixls · reddit · 2026-08-20

Testing Qwen 3.8 27B on an AMD R9700 with ROCm, the author observed that using f16 precision for KV Cache yields more careful and detailed structured/free outputs compared to q80, with more accurate reasoning and better memory retention after 120k context tokens. The author suggests that negative reviews may stem from using lower precision KV Cache like q40. A detailed llama.cpp configuration is provided.

Original post →

More from Models

Models channel →