Black Forest Labs launches Flux 3 with native audio video generation up to 20 seconds
The Decoder · rss · 2026-07-24
Black Forest Labs has released Flux 3, a multimodal foundation model trained on images, video, and audio.
Its headline feature is native sound generation for video, with clips up to 20 seconds long. The company says its own tests place it slightly ahead of Seedance 2.0, though independent benchmarks are not yet available. Black Forest Labs also says it ultimately wants to build a world model and is already testing Flux 3 on robotics tasks.
More from Embodied
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11
- SyncWorld: In-Context Robot World Model Simulates Unseen Views and Embodiments Zero-Shot — ChongZzZhang · 2026-09-11
- A 3D Pose Dataset for Dogs Released — ducha_aiki · 2026-09-11
- Swaayatt demos autonomous driving at 52 km/h on mountain roads, self-recovers after skid — sanjeevs_iitr · 2026-09-11
- AUAR's MicroFactory brings a deployable robotic wood-panel factory to the construction site — lukas_m_ziegler · 2026-09-11
- Musk: Cybercab certified at 165 Wh/mi, the most efficient production EV ever — elonmusk · 2026-09-11