Latency cut 2.7s to 0.4s on GLM via scheduling, no model or hardware changes

dbreunig · x · 2026-10-09

LunaRoute's blog argues that judging inference endpoints purely by token price is misleading: the same open-weight model can feel vastly different depending on quantization, checkpoints, caching, hardware, and scheduling.

Case in point: while optimizing their stack serving GLM-5.3-Vision-NVFP4 on an 8×B200 cluster, they cut interactive latency from 2.7s to 0.4s — a 7x improvement — by only changing how batch/background requests compete with interactive ones under load. No weight, model, or hardware changes; the tradeoff was slower background tasks.

The takeaway: "the model is not the product." As more providers serve identical open weights, the full inference stack increasingly determines the user experience.

Original post →

More from Infra

Infra channel →