FastVideo Releases 4-Step Audio-Video Model Requiring 4x B200 GPUs
NerdyRodent · x · 2026-08-29
FastVideo released the FastH3 Preview v1 checkpoint based on MiniMax H3, capable of generating synchronized video and audio in just 4 transformer forwards. The model was distilled using Data-Free DMD2 and VSA-H3 at 90% sparsity. The default setup requires 4 B200 GPUs, with specific adjustments needed for other multi-GPU CUDA systems. It currently supports text-to-audio-video generation, though fine details and audio may slightly trail the base model.
More from Multimodal
- Free 2x speed for MiniMax H3 in ComfyUI? — alisitskii · 2026-08-29
- Seedance 2.5 Demo: Audio-Driven Lip Sync & Motion — justin_hart · 2026-08-29
- Model Self-Portrait Experiment: Terra Human-like, Sol Alien, Claude Abstract — repligate · 2026-08-29
- MiniMax Music 3 produces silent or garbled audio beyond 2 minutes — God_Hand_9764 · 2026-08-29
- Turn tweets into sitcom scenes using Grok Bot and fal — altryne · 2026-08-29
- MiniMax H3: BF16 vs INT8 Convrot with Sage Attention 2.2 — Glittering-Cold-2981 · 2026-08-29