MiniMax-H3 on AMD MI355X: 26.7x faster inference, 5s video in 1.3s with prompt steering
junyanz89 · x · 2026-09-24
Nunchux runs MiniMax's MiniMax-H3 video generation model on AMD MI355X, claiming up to 26.7x faster inference than SGLang on 8 GPUs — 5 seconds of video generated in 1.3 seconds. Streaming generation lets users change the prompt while the video plays, steering what happens next. MiniMax amplified the news, calling it model optimization meeting AMD inference engineering. Free access to MiniMax-H3 is coming via a waitlist.
Related event: Nunchux speeds MiniMax-H3 on AMD MI355X by up to 26.7x(2 posts)→
More from Infra
- Alibaba's Banma ships AutoOmni 2.0: 3B-active edge model nears 10x-larger cloud models — 机器之心 · 2026-09-24
- Huang and Musk: China reaches advanced lithography in 2-3 years; chip bans only buy time — beffjezos · 2026-09-24
- MacBook Pro M5 Max runs AMD Radeon AI PRO R9700 over Thunderbolt 5 for local LLMs — TheOriginalG2 · 2026-09-24
- A real-time voice agent with zero US servers: Gladia STT, Gemma 4 on Scaleway, KugelAudio TTS — tobowers · 2026-09-24
- Running MiniCPM5-2B as a local agent on M4 16GB: trimming, tool-call adapter and benchmarks — 面壁智能 · 2026-09-24
- WSJ: The AI build-out is becoming the biggest economic bet in U.S. history — GeneReddit123 · 2026-09-24