Black Forest Labs launches Flux 3 with native audio video generation up to 20 seconds

The Decoder · rss · 2026-07-24

Black Forest Labs has released Flux 3, a multimodal foundation model trained on images, video, and audio.

Its headline feature is native sound generation for video, with clips up to 20 seconds long. The company says its own tests place it slightly ahead of Seedance 2.0, though independent benchmarks are not yet available. Black Forest Labs also says it ultimately wants to build a world model and is already testing Flux 3 on robotics tasks.

Original post →

More from Embodied

Embodied channel →