fal launches H3 Max Lip Sync: photo + audio to lip-synced video in 11 seconds
jfischoff · x · 2026-09-19
fal has launched H3 Max Lip Sync: upload one photo and one audio clip, and get a lip-synced video in seconds with natural expressions, in any language.
- Ranks #1 in fal's internal evals for both quality and speed, with a median generation time of just 11 seconds
- Built on top of fal's H3 Max architecture
- Developer isidentical says the team used diffusion RL on new verifiable tasks, and lip-sync turned out to be a surprisingly good fit — yielding what they claim is the highest-quality, fastest, and cheapest lip-sync model available
Related event: fal Launches H3 Max Lip Sync with Near-Instant Video Generation(3 posts)→
More from Multimodal
- Solo creator makes 30-minute AI sci-fi film using 1,321 generations — GlideTop · 2026-09-19
- Grok Imagine Users Made a Hollywood-Style Odyssey Scene in 9 Days for $2,677 — NicoVerderosa · 2026-09-19
- An Italian 80s/90s Canzone-Style Music LoRA Shared on Reddit — -becausereasons- · 2026-09-19
- Invideo Editor Adds Agent-Powered Sound Design for Video — azed_ai · 2026-09-19
- Upgrading from GTX 1070 Ti to RTX 5070 Ti makes Comfy video workflows 10x faster — Nesachi1 · 2026-09-19
- Redditor shares AI hunting short film made with Seedance 2 — Ok-Nerve941 · 2026-09-19