IBM runs 753B MoE on 544 H100s with llm-d, serving 3,000 concurrent coding agents at 5-10x lower cost

dl_weekly · x · 2026-09-16

IBM Research and Red Hat benchmarked llm-d serving GLM-5.2, a 753B-parameter open-weight MoE (39B active), on 544 NVIDIA H100s for 3,000 concurrent coding agents, with 85.2% of input tokens served from cache — at 5-10x lower cost than commercial APIs. The post details why agentic workloads (massive context reuse, bursty parallel sub-agents) stress inference infra, and shows enterprises can run today's agentic traffic on existing H100 fleets.

Original post →

More from coding & agent

coding & agent channel →