Running 260k Tokens on RTX 6000 via KV Quantization

AIFlow_ML · x · 2026-08-28

By using Qwen 3.8 Flash with NVFP4 and ngrams on a single RTX PRO 6000, it's possible to allocate approximately 260k tokens of KV cache in FP8. To handle multiple subagents smoothly, a KV cache quantization strategy is proposed: limiting Value to NVFP4 while keeping Key in FP8 to minimize negative impact on Softmax performance.

Original post →

More from Infra

Infra channel →