fal makes open source video model 35x faster, pushing GPUs to 70-80% of ceiling for realtime streaming
isidentical · x · 2026-09-18
In an a16z interview, fal's Gorkem Yurtseven and Batuhan Taskaya explain how they rebuilt MiniMax's open source H3 video model for faster-than-real-time generation: cutting inference steps, rewriting stage-level code, and pushing GPU utilization from the usual 30-40% to 70-80% with no quality loss — a 35x speedup.
Key points:
- An engineer live-streamed continuous H3 Max generations from a laptop on Twitch, using prompt tricks to keep a coherent story going
- fal says it's the only model capable of up to 60 minutes of continuous, action-controlled video; they capped streams at an hour, but it could run indefinitely
- The ML team noted no video model had ever generated five seconds of video in under five seconds before
- With speed and cost largely solved, the competitive frontier has shifted to quality and prompt adherence; Hollywood, not a customer a year ago, is now reaching out
Related event: fal Speeds Up MiniMax Open-Source Video Model 35x(4 posts)→
More from Infra
- Brad Gerstner at All-In Summit: who pays for AI CapEx, the gigawatt gap and semis eating the Nasdaq — DavidSacks · 2026-09-18
- Third-party audit reproduces Gensyn open-1b training step bit-for-bit — benfielding · 2026-09-18
- Anthropic Open-Sources Claude-Written GPU Optimizations Speeding 30+ Biomolecular Models ~4x — ResultBackground2450 · 2026-09-18
- Spotify: 777M users, 11-12M requests/sec — how AI changed its quality playbook — rseroter · 2026-09-18
- 605 new Linux kernel CVEs disclosed in one day, on top of 276 the day before — jedisct1 · 2026-09-18
- DeepSeek V4.1 Flash hits 532 tokens/s on Inco, fastest output on Artificial Analysis — songhan_mit · 2026-09-18