NVIDIA says it sped up MiniMax H3 video generation 3.95× in just 4.5 hours
songhan_mit · x · 2026-08-04
NVIDIA says it sped up MiniMax H3’s omni-modal video generation by 3.95× in 4.5 hours of inference-time optimization.
- MiniMax H3 is described as a 33B dense omni-modal generation model that takes text, images, video, and audio as a unified context and outputs short video clips with native stereo audio.
- The Sol Video Inference Engine got the speedup on 8× GB200, reaching 1344×768 at 24 fps and 124 frames.
- The optimization used kernel fusion, graph capture, cross-step caching, and sparse attention, with no distillation, fine-tuning, LoRA, or offline calibration.
- The blog frames this as a day-one deployment recipe for open-weight video models.
More from Infra
- Google Cloud says GKE Agent Sandbox lifts agent density from 61 to 274 per node — rseroter · 2026-08-04
- On Semiconductor beats earnings and sees AI data-center revenue more than doubling in 2026 — Polymarket · 2026-08-04
- Nuclear Startup Valar Raises $1B Led by Sequoia to Scale Reactors — kleffew94 · 2026-08-04
- AI API revenue still trails hyperscaler capex by a wide margin in 2025 chart — SurpriseDog9000 · 2026-08-04
- U.S. heartland backlash grows as AI data centers reshape local communities — altryne · 2026-08-04
- Semiconductors and data centers are being built far slower than AI demand — robleclerc · 2026-08-04