Qwen Model TurboQuant KV Cache Causes Multi-Turn Memory Loss

QuixiAI · x · 2026-08-29

A user experienced memory loss in multi-turn conversations with Qwen3.8-Flash-Next. Investigation revealed the issue was caused by the TurboQuant quantized KV Cache. Reverting the KV Cache to BF16 format resolved the problem.

Original post →

More from Models

Models channel →