MiniMax-H3 on 8×H200: 1.95× Lossless Speedup, Up to 6.24×
ying11231 · x · 2026-08-28
LMSys published a blog post detailing performance optimizations for MiniMax-H3 on an 8×H200 cluster. With fixed prompts, seeds, and resolution, the team achieved significant speedups using three layers of technology: fused kernels, Cache-DiT step reuse, and SubBlock sparse attention.
- Dense lossless path: 1.85–1.95× faster than Diffusers with no approximation.
- Max speed mode: Up to 6.24× faster using step reuse + sparse attention, though SSIM drops to 0.76–0.91.
- Presets: Conservative preset hits 2.99× at high quality (0.90–0.98 SSIM); speed preset reaches 4.90–5.93× at lower quality.
The project was a collaboration with the Cache-DiT team, Ant Group, and Nvidia.
More from Infra
- US Chip Security Act aims to verify location of high-end AI chips — peterwildeford · 2026-08-28
- GMI Router Test: 3 Tasks Split Across 3 Models, Saving $0.02-0.03 Per Prompt — Shruti_0810 · 2026-08-28
- Architect Labs claims its AI designed a chip in 2 weeks, 3.4x Jetson perf/watt — mark_k · 2026-08-28
- GMI Router Introduces KV-Cache-Aware Model Routing — anthara_ai · 2026-08-28
- Pause cloud GPUs, resume later without losing ComfyUI setup — deployonaquanode · 2026-08-28
- Guide: Deploy Agent Systems to AWS ECS with Terraform and GitHub Actions — kmeanskaran · 2026-08-28