Cloudflare Details Inference Optimizations for Running Kimi and GLM at Scale

michellechen · x · 2026-08-03

Cloudflare published a technical blog sharing underlying optimization experiences for running trillion-parameter models like Kimi and GLM at scale on their Workers AI platform.

To solve memory constraints for these large, long-context MoE models, the team applied three main techniques:

Additionally, all experiments and production traffic run and benchmark on the open-source inference framework SGLang, which the team found offers the best performance on the market.

Original post →

More from Infra

Infra channel →