Intent Lab says it sped up GLM 5.2 inference 6.3x on 2× Grace Blackwell
jiayq · x · 2026-07-29
Intent Lab says it re-engineered the TRT-LLM inference stack for GLM 5.2 and claims it is now the fastest engine it knows for that model.
- On 2× Grace Blackwell nodes, output speed improved from 102 tok/s with stock TensorRT-LLM to 647 tok/s with the optimized runtime plus DSpark.
- That is a 6.3× speedup.
- The post emphasizes that the fleet implemented the optimization end to end, not just at the model level.
Related event: Intent Lab Introduces 'fleet' and Massively Accelerates GLM 5.2 Inference(2 posts)→
More from Infra
- China’s semiconductor push adds immersion DUV and near-Micron DRAM capacity — Distinct-Question-16 · 2026-07-29
- Low-VRAM LoRA training now fits DeepSeek-V4-Flash into 90 GiB VRAM — woct0rdho · 2026-07-29
- YC startup hwintelligence launches Wave, a waveform debugger with a verification agent — ycombinator · 2026-07-29
- Broadcom says AI could cut exploit windows from weeks to hours — therealdanvega · 2026-07-29
- camelAI moved its agent from VMs to a Cloudflare Durable Object — irvinebroque · 2026-07-29
- LLM Inference Handbook collects deployment, GPU, and optimization guidance for production teams — carrycooldude · 2026-07-29