VRAM Comparison of KV Cache Quantization in Qwen3.6
token---- · reddit · 2026-07-19
The post discusses the **KV cache quantization** and VRAM usage of **Qwen3.6-35B-A3B**. The core question is whether quantizing the KV cache below **Q8** to save memory is worth it, and what the performance/quality trade-off looks like. The accompanying chart shows **VRAM usage curves** for different KV configurations across various **context lengths**. As the context grows, the VRAM difference is amplified rapidly. The unquantized approach has the highest footprint, while more aggressive quantization combinations significantly reduce memory pressure, making them ideal for long-context inference scenarios. This type of content leans more towards inference deployment and memory optimization rather than pure model capability discussions.
More from Infra
- Emad Mostaque says Kimi K3 inference costs could fall 10x to 50x soon — rohanpaul_ai · 2026-07-21
- TokenPrint turns Qwen inference into a DevTools-style visual debugger — Rich-Fruit-326 · 2026-07-21
- A broken agent router burned 30.2M tokens in 3.5 hours on Claude Code — RileyRalmuto · 2026-07-21
- Huawei's Atlas 950 SuperPoD Scales to 500,000 Chips with Unified Architecture — pstAsiatech · 2026-07-21
- South Korea's exports jump 50% in early July on the AI chip boom — Polymarket · 2026-07-21
- Bittensor boosters argue decentralized training can offset severalfold compute gaps — markjeffrey · 2026-07-21