After GLM goes live, the next test is DSv4 Flash on 8×H100

TheZachMueller · x · 2026-07-22

After deploying GLM, the author is trying to push DSv4 Flash onto 8×H100s

The post says GLM is now deployed and fast enough to use, and the next experiment is to deploy DSv4 Flash on an 8×H100 setup to see how much more throughput can be squeezed out.

The author mentions a specific parallelism configuration in mind — DP4, EP8, TP2, plus FP8 KV cache — which makes this a practical note on model-serving optimization rather than a product announcement.

Original post →

More from Infra

Infra channel →