Nunchux runs MiniMax-H3 on AMD MI355X with up to 26.7x faster inference
junyanz89 · x · 2026-09-24
- Inference startup Nunchux ported MiniMax's MiniMax-H3 to AMD MI355X GPUs, claiming up to 26.7x speedup over SGLang on 8 GPUs — a 5-second video generated in 1.3 seconds.
- Streaming generation lets users change the prompt while the video plays, steering what happens next.
- Nunchux pitches 10x cost savings and 99.9% SLA; free access to MiniMax-H3 is coming soon via waitlist.
More from Infra
- NVIDIA's DGX Spark Appears Unavailable, May Never Return to Sale — GabGarrett · 2026-09-24
- Not every AI task needs an LLM: 'decide' may become a standard model call — bigdata · 2026-09-24
- Prime Intellect launches Prime Sandboxes: MicroVM sandboxes purpose-built for RL training — xeophon · 2026-09-24
- M5 Ultra 96GB local inference: 3,200 tok/s aggregate prefill over 112M tokens with custom MLX server — Every-Fortune-3151 · 2026-09-24
- 3D DRAM slated for 2029: bandwidth of 3x HBM5E at sub-1pJ/bit — zephyr_z9 · 2026-09-24
- Brain vs H100 physical ledger: 80B transistors against 86B neurons — yaroslavvb · 2026-09-24