Tencent's STEPQuant: 6-bit recurrent states match FP32 with 68.7% less memory
_akhaliq · x · 2026-10-09
Tencent researchers released STEPQuant on Hugging Face, a spatial-temporal post-training quantization framework for Delta-rule recurrent states. It matches FP32 accuracy with 6-bit states while cutting total serving memory by up to 68.7% — a substantial deployment-cost win for linear-attention/recurrent-state models.
Related event: Tencent Open-Sources STEPQuant, Cutting Inference Memory by 68.7%(2 posts)→
More from Infra
- Bain projects 38.6M GPU and custom silicon shipments by 2030, 10x 2023's 3.9M — Beth_Kindig · 2026-10-09
- AWS reference architecture: multi-team GPU cluster sharing on SageMaker HyperPod — AWS ML Blog · 2026-10-09
- Mistral slammed for training open models on datacenters powered ~70% by coal — wavefnx · 2026-10-09
- Why do we resend the whole conversation every turn? Server-side KV slots proposal sparks debate — Vasili_Sk · 2026-10-09
- NVIDIA's NeMo-DCR cuts 1T-model weight sync from 87.5 min to 150s, 12-40x faster checkpoint transfer — dair_ai · 2026-10-09
- Zyphra Speeds MoE Expert Routing Communication 2.63x on AMD MI300X GPUs — QuentinAnthon15 · 2026-10-09