Run GLM 5.2 Models at Scale for Low Cost

Aizkmusic · x · 2026-07-17

The development team demonstrated how to run the GLM 5.2 model at scale at an extremely low cost (less than $0.1 per million tokens). By using 4 AMD MI325X GPUs, they achieved a generation speed of 1482 tokens/s, making the cost 3 times lower than B300 and 10 times lower than Opus.

Original post →

More from Infra

Infra channel →