MiniMax-H3 on 8×H200: 1.95× Lossless Speedup, Up to 6.24×

ying11231 · x · 2026-08-28

LMSys published a blog post detailing performance optimizations for MiniMax-H3 on an 8×H200 cluster. With fixed prompts, seeds, and resolution, the team achieved significant speedups using three layers of technology: fused kernels, Cache-DiT step reuse, and SubBlock sparse attention.

The project was a collaboration with the Cache-DiT team, Ant Group, and Nvidia.

Original post →

More from Infra

Infra channel →