FastVideo Releases 4-Step Audio-Video Model Requiring 4x B200 GPUs

NerdyRodent · x · 2026-08-29

FastVideo released the FastH3 Preview v1 checkpoint based on MiniMax H3, capable of generating synchronized video and audio in just 4 transformer forwards. The model was distilled using Data-Free DMD2 and VSA-H3 at 90% sparsity. The default setup requires 4 B200 GPUs, with specific adjustments needed for other multi-GPU CUDA systems. It currently supports text-to-audio-video generation, though fine details and audio may slightly trail the base model.

Related event: MiniMax open-sources H3 video model; FastH3 v1 achieves near-real-time generation(9 posts)→

Original post →

More from Multimodal

Multimodal channel →