Black Forest Labs teases FLUX.3 as an open-weights multimodal backbone

multimodalart · x · 2026-07-24

Black Forest Labs’ FLUX.3 [dev] is teased as an open-weights multimodal backbone for content creation and action prediction.

According to the post and the accompanying slide, the model is positioned for image, video, audio, and robotic action prediction, with the author highlighting the self-flow architecture as a notable evolution of flow matching. The message is that BFL is continuing to innovate at the architecture level, not just by scaling model size.

Related event: Black Forest Labs Teases Multimodal FLUX.3 Model(3 posts)→

Original post →

More from Multimodal

Multimodal channel →