Cloudflare's LLM Serving Math: FP8 and INT4 Quantization Double Concurrency

arpit_bhayani · x · 2026-08-05

The author breaks down Cloudflare's blog on serving Kimi K2.6 and GLM 5.2 at scale, noting that KV cache typically exhausts GPU memory before model weights do.

Core Optimization Strategies:

Benchmarks confirm negligible quality loss: FP8 KV cache scores within a point of BF16, and INT4 weights stay within 0.8 points of FP8.

Related event: Cloudflare Details VRAM Optimization for Large-Scale Kimi and GLM Deployment(4 posts)→

Original post →

More from Infra

Infra channel →