FLUX 3 unifies image, video, audio, and action prediction in one model

Eric520CC · x · 2026-07-24

FLUX 3 adds image, video, audio, and action prediction in one model

Black Forest Labs says FLUX 3 is a unified multimodal architecture spanning image, video, audio, and action prediction. The company says creations are more lifelike across styles, and FLUX 3 Video is now in early access.

The announcement also says the same unified architecture can be extended to predict actions for robotics, and points to work with Mimic and Audi as part of that direction.

Related event: Black Forest Labs Releases FLUX 3 Unified Multimodal Model(4 posts)→

Original post →

More from Embodied

Embodied channel →