IBM runs 753B MoE on 544 H100s with llm-d, serving 3,000 concurrent coding agents at 5-10x lower cost
dl_weekly · x · 2026-09-16
IBM Research and Red Hat benchmarked llm-d serving GLM-5.2, a 753B-parameter open-weight MoE (39B active), on 544 NVIDIA H100s for 3,000 concurrent coding agents, with 85.2% of input tokens served from cache — at 5-10x lower cost than commercial APIs. The post details why agentic workloads (massive context reuse, bursty parallel sub-agents) stress inference infra, and shows enterprises can run today's agentic traffic on existing H100 fleets.
More from coding & agent
- Radio launches: a shared chat room where agents from different providers talk directly — rohanpaul_ai · 2026-09-16
- Celesto open-sources disposable full macOS desktops for AI agents on Apple Silicon — aniketmaurya · 2026-09-16
- AI connector value lies in secrets store and personal context, not payments — jeff_weinstein · 2026-09-16
- Stripe's model-run shop bench: 5 of 7 working stores built by Claude — bcherny · 2026-09-16
- New hire ships five projects in three weeks by syncing team context with /hq-sync — jacob_posel · 2026-09-16
- Anthropic previews Model Hardware Standard to let AI agents run lab instruments — ivan_bezdomny · 2026-09-16