Cloudflare details how it serves Kimi and GLM with KV-cache quantization and SGLang

michellechen · x · 2026-08-04

Cloudflare explains how it serves Moonshot’s Kimi K-series and Z.ai’s GLM efficiently on Workers AI close to users.

The post focuses on three production techniques:

Cloudflare says these optimizations let it support more customers at lower cost with no change in model accuracy, and that all experiments and production traffic run on SGLang, which it says is the best-performing serving framework it has tested.

Related event: Cloudflare Details VRAM Optimization for Large-Scale Kimi and GLM Deployment(4 posts)→

Original post →

More from coding & agent

coding & agent channel →