Cloudflare's Memory Optimization for Serving Kimi and GLM at Scale
arpit_bhayani · x · 2026-08-05
Cloudflare shared their engineering practices for running large models like Kimi K2.6 and GLM 5.2 at scale. For long-context models, the KV cache, rather than model weights, is typically the first to exhaust GPU memory.
To increase concurrency and reduce costs, Cloudflare applied three core optimizations:
- KV Cache Quantization: Reducing Kimi K2.6's cache from BF16 to FP8 precision, instantly halving its memory footprint.
- Weight Compression: Further compressing model weights to free up space.
- Cache Protection: Ensuring cache safety and stability on shared hardware with increased request loads.
For the underlying inference framework, the team uses and actively contributes to the open-source framework SGLang, citing it as the best performing solution currently on the market.
More from Infra
- RestoreKV: Recovering Performance Under Aggressive KV Cache Eviction — Changwoo Baek · 2026-08-05
- Samsung and SK Hynix Evaluate Chinese Chipmaking Equipment to Hedge Export Controls — pstAsiatech · 2026-08-05
- Anthropic Confirms In-House Chip Team to Custom Hardware for Claude — kimmonismus · 2026-08-05
- celld: An Open-Source Self-Hosted Distributed DO Implementation — jakedahn · 2026-08-05
- NVIDIA DGX Spark Price Surges to 8,000 Euros in Europe — Afraid-Yoghurt6731 · 2026-08-05
- Google Joins HBF Consortium: TPU Could Get 10x HBM Capacity at Lower Cost — BenBajarin · 2026-08-05