a16z podcast: how fal's H3 Max turns MiniMax's video model near real-time
a16z · youtube · 2026-09-18
a16z GP Jennifer Li talks with fal co-founder Gorkem Yurtseven and Head of Engineering Batuhan Taskaya about generative video reaching real time.
- H3 Max is fal's post-trained version of MiniMax's open-weight video model; combining post-training with systems and hardware optimization slashed generation latency while keeping quality.
- That speed unlocked continuous video experiments: streams that remember previous scenes and respond to new directions live — demoed on Twitch the day after launch.
- The team argues the next bottleneck is control, not speed: camera movement, lighting, characters, motion, lip sync — pros need predictable tools, not one-shot prompt-to-video.
- Also covers inference economics, chip footprint, streaming experiences, and implications for creator workflows and the creator economy.
Full timestamped outline in the original video.
More from Infra
- fal made an open-source video model 35x faster — and Hollywood became its fastest-growing segment — a16z · 2026-09-18
- AMD plans ~10% price hike across GPUs, chipsets, and possibly CPUs — FullstackSensei · 2026-09-18
- Dev Builds 'Local ChatGPT in a Box' on One RTX 5090, Sharing Every Workaround Along the Way — valdev · 2026-09-18
- King Charles Meets OpenAI, Anthropic, DeepMind and Nvidia Execs on AI Safety — eyishazyer · 2026-09-18
- PlanetScale's TIN beats Postgres GIN full-text search: 212ms vs 288s at p99 — DanielLockyer · 2026-09-18
- NVIDIA shows 100x faster scikit-learn spectral clustering with cuML — NVIDIA Developer · 2026-09-18