fal made an open-source video model 35x faster — and Hollywood became its fastest-growing segment
a16z · x · 2026-09-18
In an a16z interview, fal's Gorkem Yurtseven and Batuhan Taskaya explain how they rebuilt MiniMax's open-source H3 video model: cutting generation steps and rewriting the code under each stage pushed GPUs from their usual 30-40% utilization to 70-80% of theoretical ceiling — 35x faster with no quality loss. Key takeaways: video now generates faster than you can film it, so price and speed are no longer the differentiators — quality and prompt adherence are. Hollywood wasn't a customer a year ago and is now fal's fastest-growing segment, using it for shot extension, camera moves and lighting tweaks that land 80-90% of the time; fal is chasing 99.9%.
More from Infra
- Google and Speakeasy open-source their entire OpenAPI SDK generation suite — _philschmid · 2026-09-18
- Dev porting $65/1B-token model to CUDA Rust to build and optimize a custom inference stack — idanbeck · 2026-09-18
- Terafab isn't a fab, it's a city: collapsing chipmaking's most fragmented supply chain — JOBhakdi · 2026-09-18
- Ben Bajarin argues compute is becoming fungible as hyperscalers redesign data center economics — BenBajarin · 2026-09-18
- LLM Studio: free Android app runs LLMs offline, generates images and turns your phone into an AI server — Miserable_Bird5822 · 2026-09-18
- AMD plans ~10% price hike across GPUs, chipsets, and possibly CPUs — FullstackSensei · 2026-09-18