Black Forest Labs launches Flux 3 with native audio video generation up to 20 seconds
The Decoder · rss · 2026-07-24
Black Forest Labs has released Flux 3, a multimodal foundation model trained on images, video, and audio.
Its headline feature is native sound generation for video, with clips up to 20 seconds long. The company says its own tests place it slightly ahead of Seedance 2.0, though independent benchmarks are not yet available. Black Forest Labs also says it ultimately wants to build a world model and is already testing Flux 3 on robotics tasks.
More from Embodied
- BeingBeyond’s Being-M0.7 learns humanoid motion from video and beats prior methods on Unitree G1 — jiqizhixin · 2026-07-24
- SN44 launches a card-grading challenge that scores five visual quality signals — bittingthembits · 2026-07-24
- New 3D foundation model paper uses Riemannian flow matching to stay on the manifold — kwangmoo_yi · 2026-07-24
- AMD’s next Gorgon Halo platform will raise unified memory to 192GB — ryanshrout · 2026-07-24
- AMD’s Ryzen AI Halo targets local AI apps with 128GB unified memory — ryanshrout · 2026-07-24
- Robocurve aims to benchmark robots on real-world physical tasks — garrytan · 2026-07-24