Qwen Model TurboQuant KV Cache Causes Multi-Turn Memory Loss
QuixiAI · x · 2026-08-29
A user experienced memory loss in multi-turn conversations with Qwen3.8-Flash-Next. Investigation revealed the issue was caused by the TurboQuant quantized KV Cache. Reverting the KV Cache to BF16 format resolved the problem.
More from Models
- Is Qwen 3.8 27B at Q2 quantization still usable? A 16GB owner asks — Effective_Head_5020 · 2026-08-29
- User Reports Opus 5.1 Fixes Response Style, Ditches Technobabble — daniel_mac8 · 2026-08-29
- Users Report Opus 5 Struggles with Instruction Following, Ignores Negative Constraints — TheOnlyVibemaster · 2026-08-29
- Is it normal to spend $100 in a few hours on GLM 5.3 API? — BLUECOW009 · 2026-08-29
- MiniMax H3 Raises Shape Mismatch Error with Audio Reference Input — Ok-Flatworm5070 · 2026-08-29
- Co-Scientist evaluation: Severe hallucinations drop to 4%, fabrication to 0% — SRSchmidgall · 2026-08-29