fal rebuilt MiniMax's open-source H3 video model to run 35x faster, pushing GPUs to 70-80% of ceiling
gorkem · x · 2026-09-18
- In an a16z interview, fal's Gorkem Yurtseven and Batuhan Taskaya explain how they rebuilt inference for MiniMax's open-source H3 video model: fewer generation steps and rewritten stage code pushed GPU utilization from the typical 30-40% to 70-80% of theoretical ceiling — a 35x speedup with no quality loss. Generation is now faster than filming, and user demand has shifted from speed/cost to quality and prompt adherence.
- Pro workflow: VFX artists render a low-res scene in Blender and feed it to the AI model for near-100% controllability; a week after H3 Max launched, an entirely new pipeline emerged — using an LLM (Astra) to build the Blender scene, then passing it to H3 Max.
More from Infra
- Hyperbolic hires quant researchers to build GPU compute as a tradable asset class — YiMaTweets · 2026-09-18
- Crusoe raises $3.9B Series F at $30.9B valuation to fuel AI energy buildout — beffjezos · 2026-09-18
- Periodic Labs details its stack: 4.1x Megatron throughput, frontier-beating science models on 1,300 H200s — hsu_byron · 2026-09-18
- Bonsai quant hits 50 tok/s at 128k context on a 24GB card, letting users run two sessions at once — julianharris · 2026-09-18
- Jeff Dean: a handful of workloads will dominate world compute, 'crying out' for specialized silicon — AccBalanced · 2026-09-18
- Silicon Data chart shows Nvidia GPUs retaining value well above depreciation schedules — AccBalanced · 2026-09-18