Upgrading inference engine boosts decode by 43% on GH200

colinmcnamara · x · 2026-08-29

After A/B testing a model on a GH200, the author upgraded the inference engine for an unrelated reason. This yielded a 43% increase in decode throughput and nearly 4x the KV cache pool with the same model and hardware.

Key Takeaways:

Original post →

More from Infra

Infra channel →