Running 260k Tokens on RTX 6000 via KV Quantization
AIFlow_ML · x · 2026-08-28
By using Qwen 3.8 Flash with NVFP4 and ngrams on a single RTX PRO 6000, it's possible to allocate approximately 260k tokens of KV cache in FP8. To handle multiple subagents smoothly, a KV cache quantization strategy is proposed: limiting Value to NVFP4 while keeping Key in FP8 to minimize negative impact on Softmax performance.
More from Infra
- Australia Datacentres Use 3% Power, Set to Hit 13% by 2035 — nordicinst · 2026-08-28
- Cloudflare saved 100TB of memory with 5 changes to 1.1.1.1's DNS cache — ritakozlov · 2026-08-28
- GPT price cuts trigger 13.8x usage surge — scaling01 · 2026-08-28
- Dual-GPU on AM5 delivers 0.1GB/s instead of 8GB/s: Promontory bridge blamed — Ed-2-Zero-9 · 2026-08-28
- DwarfStar Adds GLM 5.3 Flash Support with Q2/Q4 on MacBook — antirez · 2026-08-28
- After NVIDIA's llama.cpp acquisition, are used V100s still a safe cheap-VRAM bet? — OnlineParacosm · 2026-08-28