Qwen 3.8 27B tests show KV Cache f16 outperforms q8_0 in quality
Felixls · reddit · 2026-08-20
Testing Qwen 3.8 27B on an AMD R9700 with ROCm, the author observed that using f16 precision for KV Cache yields more careful and detailed structured/free outputs compared to q80, with more accurate reasoning and better memory retention after 120k context tokens. The author suggests that negative reviews may stem from using lower precision KV Cache like q40. A detailed llama.cpp configuration is provided.
More from Models
- Qwen3.8-27B-OBLITERATED Released: A Red-Teamed Uncensored Model — OBLITERATUS · 2026-08-20
- GPT-5.6 Ultra mode shows little coding gain over Extra High in testing — techartist_ · 2026-08-20
- Comment: GPT-5.6 Sol notorious for searching online instead of solving — zainhas · 2026-08-20
- DeepSeek V5 Suspected Testing in the Wild; Claude Code Gets Concise Mode — WorldofAI · 2026-08-20
- SpaceXAI Swaps Land for 69 Acres, Pledges $40M for Public Safety Facilities — chrisgrayson · 2026-08-20
- Users miss cold, objective AI: old prompts to cut filler talk no longer work — Wargaming123A · 2026-08-20