TurboQuant-GPU: Compresses LLM KV Cache by 5x on Any NVIDIA GPU
tom_doerr · x · 2026-08-12
The open-source project TurboQuant-GPU can compress the LLM KV cache by 5.02x during inference, working on any NVIDIA GPU.
- Mechanism: Uses random orthogonal rotation to make KV cache coordinates approximately Gaussian, followed by Lloyd-Max quantization, achieving 3 bits per element with 0.98 cosine similarity.
- Tech Stack: Built on cuTile kernels with an automatic PyTorch fallback, easily installable via pip.
More from Infra
- UBS Predicts 1 ZB KV Cache by 2030, Cuts Rubin Standard VRAM to 8.3TB — zephyr_z9 · 2026-08-12
- OpenSSH 10.5 Released: AI Becomes Major Force in Bug Discovery, Accelerating Releases — jedisct1 · 2026-08-12
- hmon: A fast C++ system monitor for Linux with GPU, Docker, and more — tom_doerr · 2026-08-12
- Why A100s Stay Racked: Monetizing Stranded Datacenter Power, Not Long GPU Lifespans — BenBajarin · 2026-08-12
- Anthropic Hiring Compute Leads to Oversee $50B+ Infrastructure Capex — surmenok · 2026-08-12
- Grace Blackwell Systems Yield 260% ROIC Under New Economic Model — BenBajarin · 2026-08-12