MiniMax-H3 on AMD MI355X: 26.7x faster inference, 5s video in 1.3s with prompt steering

junyanz89 · x · 2026-09-24

Nunchux runs MiniMax's MiniMax-H3 video generation model on AMD MI355X, claiming up to 26.7x faster inference than SGLang on 8 GPUs — 5 seconds of video generated in 1.3 seconds. Streaming generation lets users change the prompt while the video plays, steering what happens next. MiniMax amplified the news, calling it model optimization meeting AMD inference engineering. Free access to MiniMax-H3 is coming via a waitlist.

Related event: Nunchux speeds MiniMax-H3 on AMD MI355X by up to 26.7x(2 posts)→

Original post →

More from Infra

Infra channel →