4x DGX Spark Cluster Achieves 44.6 tok/s on GLM-5.2 for Real Agent Workloads

EAccelerate_42 · x · 2026-08-11

A developer shared an optimized serving configuration for GLM-5.2 running on a 4-node NVIDIA DGX Spark (GB10) cluster. The author emphasized that the setup was tuned entirely around real agent coding workloads rather than synthetic benchmark screenshots.

The cluster achieves a 44.6 tok/s single-stream decode speed, a massive jump from 5.7 tok/s just two days prior. The performance breakthrough was unlocked by a previously overlooked environment variable. The full configuration and tuning details are open-sourced on GitHub.

Original post →

More from coding & agent

coding & agent channel →