NVIDIA Sol Engine speeds up MiniMax H3 video generation by 27.7x

DavidmComfort · x · 2026-09-02

The NVIDIA SANA team utilized Sol Engine to optimize the MiniMax H3 model, implementing a pipeline of a 4-step low-res draft followed by a 3-step LTX refinement pass. This reduced 10s 768p video generation latency on a single GB200 from 414s to 14.93s (27.7x speedup). The optimization replaces heavy VAE decoders with TAEH3/TAEHV, maintaining stable latents for refinement and demonstrating co-design of sampling topology with hardware kernel acceleration. This dramatic collapse in latency fundamentally shifts unit economics, allowing a single node to serve 378K videos per month at 97%+ GPU margins, moving high-fidelity AI video from async batch rendering to near-instant, interactive infrastructure.

Original post →

More from Infra

Infra channel →