WUSH-KV: Data-Adaptive Transforms for 2-bit KV-Cache Quantization Integrated into SGLang
ISTA-DASLab · hf · 2026-10-01
- KV cache memory and bandwidth costs grow with context length and batch size; ISTA-DASLab proposes WUSH-KV for low-bit KV-cache quantization.
- WUSH constructs a data-aware transform from second-order statistics of both factors in a matrix product: separate key/value transforms are built from calibration data, with the value transform folded into weights and the key transform applied after RoPE.
- Under mild assumptions the WUSH transform is near-optimal for the QuEST INT quantizer. Integrated into SGLang with OSCAR-style percentile-clipped affine quantization, WUSH-KV reduces layerwise reconstruction error, achieves the lowest end-to-end perplexity among tested transforms, and at 2-bit matches or outperforms the OSCAR transform across all models and tasks.
More from Infra
- Delip Rao: Most big-budget GPU training runs are run sub-optimally — deliprao · 2026-10-01
- FT: Tencent leases 100,000 AI chips from Oracle for $7B over 5 years — rohanpaul_ai · 2026-10-01
- Silicon microring modulators push past 200Gb/s per lane to cut AI optical I/O power — jwt0625 · 2026-10-01
- Linewise's video agent hits 31x GPU throughput at 1/15 cost on Inco inference infra — songhan_mit · 2026-10-01
- Memory stocks rally overnight in Asia as AI demand frenzy reignites — firstadopter · 2026-10-01
- Boat VMs claimed cheapest scalable VM infra, could save hundreds of thousands monthly on AI sandboxes — Scobleizer · 2026-10-01