Black Forest Labs Launches FLUX 3: A Unified Multimodal Model for Image, Video, and Audio

dl_weekly · x · 2026-08-04

Black Forest Labs announced that its latest multimodal foundation model, FLUX 3, is now available in Early Access. The model uses a unified architecture to jointly train on images, videos, and audio.

The company stated that this joint learning approach allows the model to better understand the underlying rules of the physical world, such as the causal relationship between object motion and sound. Currently, FLUX 3 can generate 20-second video clips with native audio and shows potential to expand into physical AI applications like robotic action prediction.

Original post →

More from Models

Models channel →